DevOps Manager · Ren
Jan 2025 to PresentLead a 7-person DevOps team running AWS (EKS) and Azure (AKS) platforms for every product team in the company. Replaced a patchwork of tools with one uniform delivery solution and built the AI layer the team runs on.
- Took the organization from an assortment of tools, pipelines and manual deploys to one uniform delivery solution across AWS and Azure: all infrastructure in Terraform, build and test in GitHub Actions, and deployment through Argo CD GitOps.
- Retired Octopus Deploy and TeamCity in favor of GitHub Actions and Argo CD, using an AI-assisted migration skill that converts each service's build pipeline.
- Brought all infrastructure across every AWS account and Azure subscription under Terraform, with coverage reported weekly to department heads.
- Made Argo CD the company-standard GitOps engine and retired Flux. Piloted a hub-and-spoke layout with argocd-agent and fixed a server-side-apply diff problem that was blocking delivery from the hub.
- Built new platforms on the standard pattern: a product platform across four AWS environments (self-hosted GitHub Actions runners on EKS, Datadog, private ingress, per-account secrets) and Temporal on AWS and Azure with mTLS.
- Automated Kubernetes upgrades for EKS, AKS and Argo CD (cluster discovery, version-path analysis, preflight audit, dry run) behind a browser-based upgrade command center. Migrated staging and production from RDS to Aurora with a reusable scripted cutover.
- Built an in-house AI tier 1 on-call assistant that is the first line of defense on S1 alerts. With access to FireHydrant, Jira, Confluence and the team knowledge base, it investigates the alert, researches what is going on, and gives the on-call engineer a proposed fix. The engineer decides and acts.
- Built the team brain into Claude Code as a shared plugin: architecture, runbooks and conventions load into every engineer's session, so engineers become subject-matter experts much faster and can respond to issues across every team in the company.
- Designed an agentic engineering-management layer on Claude Code: custom skills and scheduled headless agents that pull from Jira, Datadog, Outlook and Teams to run sprint planning, backlog grooming, help desk triage and routing, on-call review and weekly ops reporting.
- Enforced AI guardrails as policy-as-code: hooks shipped with the team plugin stop agents from making non-compliant commits and PRs or closing tickets without a root cause.
- Used agents with persistent memory to root-cause long-standing production issues: an N+1 call fan-out behind recurring 504s, shared-library memory leaks causing out-of-memory kills, and duplicate private DNS zones behind an outage, found by auditing DNS across every AKS cluster.
- Scoped an AI observability platform (Langfuse on dedicated AWS accounts and EKS) and set a multi-agent coding pattern where Opus orchestrates and Sonnet subagents handle routine work, to keep model cost down.
- Sponsored two engineers for senior promotion, onboarded a senior hire, and introduced sprint working agreements and a complexity-weighted help desk rotation.