Job Description
Senior Cloud Operations \& AI Automation Engineer
What if every runbook you wrote could execute itself, every root-cause analysis could draft its own findings, and every 3 AM page could be met by an autonomous agent you trained — before you even open your laptop?
We run an AI-native community and social engagement platform that powers customer relationships for the world's largest brands. A single outage here isn't an internal inconvenience — it's a headline-level event for a Fortune 100 client. That's the bar. We're hiring a senior operations engineer to hold it.
The Role
You are an experienced, hands-on cloud operator who also happens to be an AI-native builder. You carry the pager, you resolve the incident, and then you do the part most teams skip: you encode what you learned into an agent, a guardrail, or a self-healing workflow so that class of problem never requires a human again.
Your performance isn't measured by ticket velocity. It's measured by the growing surface area of production work your agents can handle autonomously — safely, within policy, with humans in the loop where it matters.
What You'll Own
- On-call \& incident command. First responder on your shift. Triage, resolve, and escalate customer-impacting events. You treat uptime like a personal commitment, not a team abstraction.
- AI agent lifecycle. Design, deploy, tune, and maintain the autonomous agents that power pre-triage, change validation, auto-healing, RCA drafting, and permanent-fix tracking. When an agent fails, you fix the capability — not just the output.
- Safe production execution. Deployments, configuration changes, and cost-optimization actions pass through quality gates with tested rollbacks. You pull the lever the instant a change drifts off-plan.
- Root-cause discipline. Investigate to the actual cause — not the symptom — then close the loop by shipping the prevention. An RCA without a landed fix is incomplete work.
- Knowledge multiplication. Document every procedure so agents can retrieve it and teammates are never starting cold. In a global async team, unwritten knowledge is lost knowledge.
What We Expect
- 5+ years in production SRE, DevOps, Platform Engineering, or Cloud Operations roles — real pager time on large-scale SaaS, not exclusively build-and-hand-off work.
- Deep AWS expertise at scale — multi-AZ, multi-account architectures, production incident management, and infrastructure automation. You've accumulated real operational scars and can narrate them in detail.
- Self-directed ownership. You run your shift like it's your company. No one assigns your priorities or audits your hours. You spot the gap, ship the fix, and raise the bar — or push back with a better plan.
- AI-native operating style. You delegate meaningful units of work to agents, evaluate their output critically, and iterate on the underlying capabilities when they fall short. Proficient with modern agentic tooling (Claude Code, Codex, Warp, custom agents) and eager to evaluate new models the day they ship.
- AWS Solutions Architect – Associate (or equivalent production depth that makes the cert redundant).
- Fluent English — clear and precise in writing and under pressure on an incident bridge.
- Shift-based availability — on-call rotations and time-zone coverage are core to the role, not an afterthought.
- OFAC-clear country of residence.
Nice to Have
- Published work, open-source contributions, or original tooling in agentic operations or AIOps.
- Experience with multi-tenant B2B SaaS — community platforms, customer-experience tools, or observability products.
- Familiarity with modern observability stacks: Grafana, Prometheus, Datadog, PagerDuty, OpsGenie.
- Azure knowledge alongside deep AWS experience.
- A documented deep-dive obsession — a hard problem you stayed with long enough to master, inside or outside your day job.
What You'll Gain
Hands-on experience building one of the industry's first truly agentic operations organizations — not reading case studies, but shipping the agents, runbooks, and guardrails that define what AI-native reliability looks like at enterprise scale. You'll leave with skills and stories the rest of the industry is still trying to understand.
Working Conditions
- Enterprise stakes, startup tempo. Fortune 100 clients and contractual SLAs paired with a small-team operating model — weekly delivery cadence, fast decisions, constant iteration.
- No artificial constraints on tooling. If the right solution requires more compute, a better model, or a tool we haven't purchased yet, we invest.
- Fully remote, global team.