We are seeking a Senior Cloud Operations & AI Automation Engineer to join our AI-native community and social engagement platform team. This role is critical in maintaining the reliability and uptime of services that support some of the world’s largest brands. You will be a hands-on cloud operator and AI builder, responsible not only for incident response but also for automating solutions that prevent recurring issues. Your work will directly impact high-profile clients, where outages are significant events, and your success will be measured by the autonomous capabilities you develop to handle production work safely and effectively.
Key Responsibilities
- Lead on-call and incident command duties as the first responder during your shift, managing triage, resolution, and escalation of customer-impacting events with a strong personal commitment to uptime.
- Design, deploy, tune, and maintain AI-driven autonomous agents that support pre-triage, change validation, auto-healing, root cause analysis (RCA) drafting, and tracking permanent fixes. When agents fail, you will address the underlying capability, not just the symptoms.
- Ensure safe production execution by managing deployments, configuration changes, and cost-optimization actions through quality gates with tested rollback procedures. You will act decisively if any change deviates from the plan.
- Conduct thorough root cause investigations to identify the actual cause of incidents and close the loop by implementing permanent fixes. An RCA without a fix is incomplete.
- Document procedures comprehensively to enable agents to retrieve knowledge and ensure teammates have access to critical information, supporting a global asynchronous team environment.
Required Qualifications
- Minimum of 5 years’ experience in production Site Reliability Engineering (SRE), DevOps, Platform Engineering, or Cloud Operations roles with hands-on pager responsibility in large-scale SaaS environments.
- Deep expertise in AWS at scale, including multi-AZ and multi-account architectures, production incident management, and infrastructure automation. You should be able to share detailed operational experiences and lessons learned.
- Strong self-direction and ownership mindset, managing your shift independently without assigned priorities or time audits. You proactively identify gaps, implement fixes, and raise operational standards or propose better plans.
- AI-native operating approach, delegating meaningful tasks to autonomous agents, critically evaluating their outputs, and iterating on capabilities when needed. Proficiency with modern agentic tools such as Claude Code, Codex, Warp, or custom agents is essential, along with eagerness to adopt new AI models promptly.
- AWS Solutions Architect – Associate certification or equivalent practical expertise that renders the certification redundant.
- Fluent English communication skills, both written and verbal, especially under pressure during incident management.
- Availability for shift-based work, including on-call rotations and time-zone coverage as a core responsibility.
- Residency in an OFAC-compliant country.
Preferred Qualifications and Benefits
- Contributions to published work, open-source projects, or original tooling in agentic operations or AIOps.
- Experience with multi-tenant B2B SaaS platforms such as community engagement, customer experience tools, or observability products.
- Familiarity with modern observability and incident management stacks like Grafana, Prometheus, Datadog, PagerDuty, and OpsGenie.
- Additional knowledge of Azure cloud services alongside deep AWS expertise.
- A demonstrated passion for deep-dive problem-solving, exemplified by mastering complex challenges within or outside your professional role.
What You’ll Gain
You will gain hands-on experience building one of the industry’s pioneering agentic operations organizations. This is a unique opportunity to create and ship the agents, runbooks, and guardrails that define AI-native reliability at enterprise scale. Your skills and insights will place you at the forefront of an evolving field that many are still striving to understand.
Working Conditions
- Operate at enterprise stakes with startup agility, supporting Fortune 100 clients under contractual SLAs while working within a small, fast-moving team that delivers weekly and iterates constantly.
- Access to unrestricted tooling investments; if a solution requires more compute, better models, or new tools, the company will provide them.
- Fully remote role within a global team, offering flexibility and diverse collaboration opportunities.