Summary

DevOps manager who replaced a patchwork of tools, pipelines and manual deploys with one uniform delivery platform across AWS and Azure: all infrastructure in Terraform, build in GitHub Actions, deployment through Argo CD. Built the AI layer the team runs on, including an in-house tier 1 on-call assistant, a shared Claude Code knowledge base, and agentic workflows with policy guardrails. Kubestronaut, 10 years in DevOps and 14 in infrastructure.

7Person DevOps team led
1Delivery standard across AWS and Azure
100%Infrastructure in Terraform
5 of 5Kubernetes certifications

Experience

DevOps Manager · Ren

Jan 2025 to Present

Lead a 7-person DevOps team running AWS (EKS) and Azure (AKS) platforms for every product team in the company. Replaced a patchwork of tools with one uniform delivery solution and built the AI layer the team runs on.

Unified delivery platform
  • Took the organization from an assortment of tools, pipelines and manual deploys to one uniform delivery solution across AWS and Azure: all infrastructure in Terraform, build and test in GitHub Actions, and deployment through Argo CD GitOps.
  • Retired Octopus Deploy and TeamCity in favor of GitHub Actions and Argo CD, using an AI-assisted migration skill that converts each service's build pipeline.
  • Brought all infrastructure across every AWS account and Azure subscription under Terraform, with coverage reported weekly to department heads.
  • Made Argo CD the company-standard GitOps engine and retired Flux. Piloted a hub-and-spoke layout with argocd-agent and fixed a server-side-apply diff problem that was blocking delivery from the hub.
  • Built new platforms on the standard pattern: a product platform across four AWS environments (self-hosted GitHub Actions runners on EKS, Datadog, private ingress, per-account secrets) and Temporal on AWS and Azure with mTLS.
  • Automated Kubernetes upgrades for EKS, AKS and Argo CD (cluster discovery, version-path analysis, preflight audit, dry run) behind a browser-based upgrade command center. Migrated staging and production from RDS to Aurora with a reusable scripted cutover.
AI operations
  • Built an in-house AI tier 1 on-call assistant that is the first line of defense on S1 alerts. With access to FireHydrant, Jira, Confluence and the team knowledge base, it investigates the alert, researches what is going on, and gives the on-call engineer a proposed fix. The engineer decides and acts.
  • Built the team brain into Claude Code as a shared plugin: architecture, runbooks and conventions load into every engineer's session, so engineers become subject-matter experts much faster and can respond to issues across every team in the company.
  • Designed an agentic engineering-management layer on Claude Code: custom skills and scheduled headless agents that pull from Jira, Datadog, Outlook and Teams to run sprint planning, backlog grooming, help desk triage and routing, on-call review and weekly ops reporting.
  • Enforced AI guardrails as policy-as-code: hooks shipped with the team plugin stop agents from making non-compliant commits and PRs or closing tickets without a root cause.
  • Used agents with persistent memory to root-cause long-standing production issues: an N+1 call fan-out behind recurring 504s, shared-library memory leaks causing out-of-memory kills, and duplicate private DNS zones behind an outage, found by auditing DNS across every AKS cluster.
  • Scoped an AI observability platform (Langfuse on dedicated AWS accounts and EKS) and set a multi-agent coding pattern where Opus orchestrates and Sonnet subagents handle routine work, to keep model cost down.
Leadership
  • Sponsored two engineers for senior promotion, onboarded a senior hire, and introduced sprint working agreements and a complexity-weighted help desk rotation.

Senior Manager of Cloud Operations · Lucidworks

May 2022 to Dec 2024
  • Led a team of 8 engineers handling both project work and interrupt-driven operations for enterprise search deployments across multiple clouds.
  • Owned development, staging and production environments for enterprise clients on Kubernetes, using Terraform, Argo CD, Helm and custom Python and Go tooling to meet a 99.99% SLA.
  • Introduced a structured incident response and root-cause-analysis process, with monitoring and alerting in Prometheus, Grafana and PagerDuty.
  • Directed the design of an internal toolbox that consolidated a dozen operational tools into one microservices web application, adopted across teams.
  • Led the high-level design of Kubernetes environments, including backup and disaster recovery for stateful workloads, and built reusable Terraform modules and Helm charts adopted as organizational standards.
  • Worked with senior leadership to align the operations roadmap with company goals.

Principal DevOps Engineer · SemanticBits, Herndon, VA

Oct 2018 to May 2022
  • Designed highly available AWS infrastructure (EC2, EKS, ECS, RDS, S3) for government healthcare applications with strict compliance requirements.
  • Ran production Kubernetes on EKS with both Fargate and managed node groups.
  • Built immutable dev, stage and prod environments with Terraform workspaces and reusable modules, and deployed redundant Azure infrastructure (AD, Key Vault, Storage, VMs, Functions) with Terraform.
  • Containerized React, Drupal and Java applications with hardened multi-stage Docker builds and security scanning in CI.
  • Automated JMeter load testing in Jenkins so every release was validated against expected concurrency before going live.

DevOps Engineer · MNX Solutions

Oct 2016 to Oct 2018
  • Built and ran AWS infrastructure (EKS, EC2, RDS, S3, API Gateway, ALB) for high-traffic e-commerce and logistics clients, including containerizing a legacy .NET 3.5 application onto EKS.
  • Created reusable CloudFormation templates and Ansible playbooks for single-command deployment of load-balanced Java stacks with Jenkins CI/CD.
  • Deployed Kubernetes with KOPS and Terraform, with Spinnaker for canary and blue-green releases, and built pipelines for PHP, Python, Go and Node.js services on GitLab, GitHub, Jenkins and CircleCI.

Senior Systems Engineer · arakÿta, Toledo, OH

Jan 2012 to Oct 2016
  • Built a private cloud on a VMware cluster behind redundant Cisco ASA firewalls, hosting a multi-tenant platform (RemoteApps, Exchange, SharePoint) for 1,000+ users.
  • Engineered a ZFS snapshot replication backup product with 30-day retention and minimal bandwidth, and a Splunk health dashboard fed by SNMP and PowerShell.
  • Led AD migration, disaster recovery and P2V projects as engineer and project manager.

Earlier career: PACS engineer at Alpha Imaging, designing, deploying and supporting medical imaging systems for hospitals and clinics.

Selected projects

Self-hosted AI operations assistant

A Slack assistant with tools across five MCP servers (Kubernetes, GitOps, observability, databases) that answers questions about live infrastructure. It hands multi-step changes to a custom Kubernetes operator that runs Claude Agent SDK jobs and waits for approval in Slack before acting.

Python, FastAPI, Go, Claude Agent SDK, MCP, Kubernetes operator, Argo CD

LLM-assisted production error triage

For a legacy .NET/IIS production estate: errors are grouped into signatures every five minutes and filtered by a rules layer kept in git. Only new or flagged signatures go to Claude for root-cause triage, capped per cycle, which keeps model spend and alert noise low.

Python, Loki, Claude, GKE, Slack

Multi-user AI development room

People and Claude edit a shared codebase in real time. The AI proposes each change and a person approves or rejects it, with automatic Haiku, Sonnet and Opus routing and per-session and per-day spend limits.

Node.js, Fastify, Socket.IO, React, isomorphic-git, Kubernetes

Legacy platform migration to GCP

Moved a national sports organization's .NET/IIS estate from managed hosting to GCP and cut over production, then root-caused a post-cutover outage to a FUSE storage mount and removed the dependency by serving files from Cloud Storage.

GCP, managed instance groups, Cloud SQL, Cloud Load Balancing, Certificate Manager

Certifications

  • Kubestronaut 2025The Linux Foundation: CKA (2024), CKAD, CKS, KCNA, KCSA (2025)
  • Professional Cloud Architect 2023Google Cloud
  • Associate Cloud Engineer 2023Google Cloud
  • DevOps Engineer, Professional 2021AWS
  • MCSE: Cloud Platform and Infrastructure 2016Microsoft
  • MCSA: Windows Server 2012 2014Microsoft

Skills

AI and automation
Claude Code (skillspluginshooks)Claude Agent SDKModel Context Protocolmulti-agent orchestrationLangfusen8nAnthropic and OpenAI APIs
GitOps and delivery
Argo CDGitHub ActionsTerraformHelmGitLab CIJenkins
Kubernetes
EKSAKSGKEkubeadmKarpenterargocd-agentcert-manager
Cloud
AWS (EKSAuroraRDSIAMCloudFrontSecrets Manager)Azure (AKSService BusFunctionsPrivate DNS)GCP (GKECloud SQLLoad Balancing)
Observability and incidents
DatadogPrometheusGrafanaLokiFireHydrantPagerDuty
Languages
PythonGoBashPowerShellSQL
Data and messaging
PostgreSQLAuroraMySQLSQL ServerRedisTemporalElasticsearch

Education

Computer Science, Owens Community College