The Site Reliability and Observability Engineer plays a critical role in maintaining the reliability, availability, scalability, and performance of essential systems and applications. This position involves close collaboration with cross-functional teams to establish comprehensive digital monitoring and tracing capabilities. The goal is to gain proactive, actionable insights into system performance, enabling early anomaly detection and reducing mean time to recovery (MTTR).
Key Responsibilities
- Collaborate with engineering, operations, and other stakeholders to understand enterprise architecture, monitoring requirements, and performance objectives.
- Identify and define key performance indicators (KPIs), diagnose issues, and proactively highlight areas for system optimization.
- Develop and implement observability frameworks, tools, and processes to enable thorough monitoring, logging, and tracing of applications and infrastructure.
- Ensure infrastructure and deployment environments are highly available, scalable, and reliable.
- Implement and manage monitoring and observability tools such as AppDynamics, DataDog, Splunk, ELK, and Sentry to provide insights into system health and performance.
- Deliver timely and accurate reports on application performance, emphasizing key trends and actionable insights.
- Work closely with digital teams to drive performance improvements through code optimization and infrastructure adjustments.
- Provide guidance and training to internal teams and end-users on best practices for Application Performance Monitoring (APM) and performance optimization.
Required Qualifications
- A minimum of 8 years of experience in IT infrastructure and applications, with 3 to 5 years of hands-on experience in observability and continuous integration.
- At least 2 years of programming experience in Java or related technologies.
- Strong knowledge of cloud infrastructure, particularly Microsoft Azure, and experience with cluster management tools like Kubernetes.
- Proven ability to communicate complex technical concepts effectively and align stakeholders across the organization.
- Bachelor’s or Master’s degree in Information Technology, Computer Science, or a related quantitative field.
Additional Responsibilities and Environment
As a member of the digital team, this role is responsible for introducing the latest tools and innovations in the observability domain. The engineer will provide expert recommendations on monitoring systems, logging frameworks, and distributed tracing platforms. Managing and delivering key KPI metrics across the enterprise architecture and performing trend analysis are also essential duties. The role demands the delivery of a proactive monitoring framework covering both infrastructure and digital experience monitoring.
Problem Solving and Decision-Making
The engineer will apply proactive approaches to monitoring challenges by leveraging existing observability tools and domain expertise. A deep understanding of application performance metrics, monitoring, and troubleshooting is required. The role involves providing expertise in problem detection, isolation, and root cause analysis (RCA) during incident management, using relevant data and artifacts from observability tools and associated systems. Decision-making authority includes recommending best practices for observability and Site Reliability Engineering (SRE) and designing and implementing monitoring solutions.
This position offers an opportunity to work at the forefront of observability and reliability engineering, contributing to the continuous improvement of digital systems and infrastructure performance.