We are seeking a Senior Lead DevOps Engineer to design, automate, and manage secure, highly available, and resilient production platforms on AWS. This role demands extensive hands-on expertise in cloud infrastructure, automation tools, container orchestration, Linux systems, and production incident management. The ideal candidate will lead DevOps architecture and implementation efforts, establish engineering best practices, support critical production environments, and mentor team members to drive operational excellence.
Key Responsibilities
- Lead the design and implementation of DevOps solutions, focusing on secure and scalable AWS environments.
- Architect and manage multi-Availability Zone and disaster recovery strategies to ensure high availability and business continuity.
- Develop and maintain Infrastructure as Code using Terraform to automate provisioning and configuration management.
- Build and manage CI/CD pipelines using Jenkins integrated with Git or Bitbucket repositories.
- Deploy and operate containerized applications using Docker and Kubernetes.
- Monitor production systems using ELK Stack, Filebeat, Grafana, and AWS CloudWatch to ensure system health and performance.
- Administer Linux servers (Ubuntu or RHEL), including configuration, patching, and troubleshooting.
- Write and maintain automation scripts in Python and Shell to streamline operational tasks.
- Configure and support web servers such as Nginx and Apache Tomcat.
- Lead production incident response, perform root cause analysis, and implement corrective actions to minimize downtime.
- Mentor and guide engineering teams on DevOps best practices and complex technical challenges.
- Support business-critical production systems while adhering to defined SLAs and operational standards.
Required Qualifications
- Over 13 years of experience in DevOps, cloud engineering, platform engineering, or production operations roles.
- Proven expertise architecting and managing AWS environments, including services like EC2, EBS, S3, VPC, ALB, NLB, EFS, IAM, Lambda, SES, and CloudWatch.
- Strong background in designing high-availability, multi-Availability Zone, and disaster recovery solutions.
- Advanced skills in Terraform and Infrastructure as Code methodologies.
- Extensive experience with CI/CD tools, particularly Jenkins, and source control systems such as Git or Bitbucket.
- Hands-on proficiency with Docker and Kubernetes container orchestration.
- Solid experience with monitoring and logging tools including ELK Stack, Filebeat, and Grafana.
- Deep knowledge of Linux system administration on Ubuntu or RHEL platforms.
- Proficient in Python and Shell scripting for automation and operational tasks.
- Experience configuring and managing Nginx and Apache Tomcat web servers.
- Strong troubleshooting, debugging, incident management, and root cause analysis capabilities.
- Demonstrated ability to lead complex DevOps initiatives and mentor engineering teams effectively.
- Experience supporting mission-critical production systems within strict SLA requirements.
Preferred Qualifications and Benefits
While specific benefits are not detailed, this role offers an opportunity to work at a senior leadership level within a dynamic environment, driving innovation and operational excellence in cloud infrastructure and DevOps practices. Candidates with a passion for mentoring and leading technical teams will find this position rewarding.
This role requires a proactive, hands-on engineer with a comprehensive understanding of cloud platforms, automation, and production support, ready to take ownership of complex systems and lead transformative DevOps initiatives.