We are seeking a skilled DevOps Engineer with 3 to 6 years of experience in DevOps, Site Reliability Engineering (SRE), cloud engineering, or infrastructure management. The ideal candidate will have hands-on experience working with major cloud platforms such as AWS, Tencent Cloud, or Azure, and possess a strong understanding of core cloud services. This role requires a proactive professional who is comfortable managing complex cloud environments, automating infrastructure, and supporting high-traffic streaming platforms.
Key Responsibilities
- Build and maintain continuous integration and continuous deployment (CI/CD) pipelines for web, mobile, backend, and data services to ensure smooth and reliable software delivery.
- Manage cloud infrastructure primarily on AWS, with additional exposure to Azure as needed, ensuring optimal performance and scalability.
- Automate infrastructure provisioning and configuration using Infrastructure as Code tools like Terraform.
- Containerize applications and manage container orchestration using Docker and Kubernetes.
- Implement comprehensive monitoring, logging, alerting, and observability practices to maintain system health and quickly identify issues.
- Support secure, reliable, and repeatable production deployments in collaboration with cross-functional teams.
- Work closely with Engineering, QA, Security, and Product teams to align infrastructure capabilities with business needs.
- Investigate and resolve incidents, performance bottlenecks, and infrastructure-related issues through root-cause analysis.
- Support backup strategies, disaster recovery plans, and ensure high availability and scalability of services.
- Continuously improve deployment speed, operational efficiency, and provide clear visibility into infrastructure costs.
Required Qualifications
- 3 to 6 years of practical experience in DevOps, SRE, cloud engineering, or infrastructure roles.
- Hands-on expertise with AWS, Tencent Cloud, or Azure, along with a solid understanding of their core cloud services.
- Strong experience with CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, or similar platforms.
- Proficiency in container technologies including Docker and Kubernetes.
- Solid knowledge of Linux systems, networking fundamentals, and monitoring tools.
- Familiarity with Infrastructure as Code tools, particularly Terraform or equivalent.
- Understanding of security best practices, Identity and Access Management (IAM), secrets management, and operational standards.
- Experience supporting high-traffic Over-The-Top (OTT) or streaming platforms.
- Knowledge of autoscaling, Content Delivery Networks (CDN), caching mechanisms, databases, and distributed systems.
- Relevant certifications in AWS, Azure, Kubernetes, or Terraform are highly desirable.
- Strong automation mindset with excellent incident response and root-cause analysis skills.
- Security-conscious approach to infrastructure management, including clear documentation and effective technical communication.
Preferred Qualifications and Benefits
- Experience working in environments that demand high availability and rapid scaling.
- Demonstrated ability to collaborate effectively across multiple teams to deliver secure and reliable infrastructure solutions.
- Commitment to continuous learning and staying updated with emerging cloud technologies and best practices.
This role offers the opportunity to work in a dynamic environment supporting cutting-edge streaming platforms, where your expertise will directly impact the reliability and scalability of critical services. If you are passionate about cloud infrastructure, automation, and operational excellence, this position provides a challenging and rewarding career path.