We are seeking a skilled DevOps Engineer with 3 to 6 years of experience in DevOps, Site Reliability Engineering (SRE), cloud engineering, or infrastructure management. The ideal candidate will have hands-on experience working with major cloud platforms such as AWS, Tencent Cloud, or Azure, and possess a strong understanding of core cloud services. This role requires a proactive professional who is well-versed in automation, infrastructure as code, container orchestration, and security best practices to support high-traffic OTT or streaming platforms.
Key Responsibilities
- Design, build, and maintain CI/CD pipelines for web, mobile, backend, and data services to ensure smooth and reliable software delivery.
- Manage cloud infrastructure primarily on AWS, with additional exposure to Azure as needed.
- Automate infrastructure provisioning and configuration using Infrastructure as Code tools like Terraform.
- Containerize applications and manage them using Docker and Kubernetes to support scalable deployments.
- Implement comprehensive monitoring, logging, alerting, and observability practices to maintain system health and performance.
- Ensure secure, reliable, and repeatable production deployments by following operational best practices.
- Collaborate closely with Engineering, QA, Security, and Product teams to align infrastructure with business needs.
- Investigate and resolve incidents, performance issues, and infrastructure bottlenecks promptly.
- Support backup strategies, disaster recovery plans, scalability, and high availability to minimize downtime.
- Continuously improve deployment speed, operational efficiency, and provide clear visibility into infrastructure costs.
Required Qualifications
- 3 to 6 years of practical experience in DevOps, SRE, cloud engineering, or infrastructure roles.
- Hands-on experience with cloud platforms such as AWS, Tencent Cloud, or Azure, with a strong grasp of core cloud services.
- Proficiency in CI/CD tools including Jenkins, GitHub Actions, GitLab CI, or similar platforms.
- Solid experience working with Docker, Kubernetes, Linux operating systems, networking concepts, and monitoring tools.
- Familiarity with Infrastructure as Code tools, preferably Terraform.
- Strong understanding of security principles, Identity and Access Management (IAM), secrets management, and operational best practices.
- Experience supporting high-traffic OTT or streaming platforms.
- Knowledge of autoscaling, Content Delivery Networks (CDN), caching mechanisms, databases, and distributed systems.
- Relevant certifications in AWS, Azure, Kubernetes, or Terraform are highly desirable.
- Demonstrated automation and reliability mindset, with excellent incident response and root cause analysis skills.
- Security-conscious approach to infrastructure management, including clear documentation and effective technical communication.
Preferred Qualifications and Benefits
- Experience with Azure cloud services is a plus.
- Candidates with certifications related to cloud platforms and container orchestration will be given preference.
- Opportunity to work in a dynamic environment supporting large-scale streaming platforms.
- Collaborative culture with cross-functional teams focused on innovation and operational excellence.
This role offers the chance to contribute to critical infrastructure that supports high-demand streaming services, emphasizing automation, security, and scalability. If you are passionate about cloud technologies and infrastructure reliability, this position provides an excellent platform to grow your career.