Lead Cloud Engineer - AI Ops - Remote

UnitedHealth GroupMinnetonka, MN
$112,700 - $193,200Remote

About The Position

The Optum Insight Engineering team is seeking a Lead AI Ops Cloud Engineer with Site Reliability Engineering (SRE) experience to design, build, and scale modern AI Ops solutions across payer and provider transformation initiatives. This is a hands-on technical leadership role requiring deep involvement in architecture and engineering while leading globally distributed teams. The role focuses on delivering, Agentic AI, and Retrieval Augmented Generation (RAG) solutions with enterprise-grade reliability, security, and responsible AI practices. You’ll enjoy the flexibility to work remotely from anywhere within the U.S. as you take on some tough challenges. For all hires in the Minneapolis or Washington, D.C. area, you will be required to work in the office a minimum of four days per week.

Requirements

  • 8+ years of overall software engineering and Site Reliability Engineering (SRE) experience in a public cloud environment (GCP, AWS, Azure)
  • 3+ years of demonstrated hands-on experience with Python and Terraform based development
  • 2+ years delivering AI/ML or Generative AI solutions in production
  • 2+ years of experience managing Kubernetes environments (EKS, AKS, GKE, or self-hosted)
  • Available to work rotating 24x7 primary and secondary on-call shifts

Nice To Haves

  • Bachelor’s degree in Computer Science, Software Engineering, or a related technical field (or 8+ years of equivalent software engineering experience in lieu of degree)
  • Hands-on experience with cloud infrastructure automation, observability tools, and SRE best practices for AI workloads
  • Experience leading globally distributed technical teams
  • Experience with vector search and enterprise search solutions
  • Proven experience building Retrieval Augmented Generation (RAG) pipelines, Agentic AI, or multi-step AI workflows
  • Proven experience building evaluation and monitoring frameworks for AI quality
  • CI/CD (GitHub Actions preferred) & Infrastructure as Code experience with Terraform
  • Proven effective communication skills with ability to explain complex technical concepts to diverse stakeholders

Responsibilities

  • Lead end-to-end design and implementation of AI Ops solutions from concept through production with an emphasis on responsible AI practices
  • Contribute to the improvement of our SRE practices; Lead incident response; Ensure high availability, scalability, and performance of cloud environments
  • Automate Infrastructure & Operations: Develop Infrastructure as Code using Terraform & GitHub Actions while adhering to best practices
  • Define and own solution architecture for RAG pipelines, agentic workflows, tool calling, orchestration, and conversational context management
  • Design, develop, and deploy AI-powered solutions to address complex business challenges across enterprise scale using RAG based solutions
  • Work closely with development and SRE teams to improve system design, advocate for reliability, and mentor fellow engineers
  • Design, code, test, and operate software using Python & Node.js
  • Leverage enterprise-approved AI tools to streamline workflows, automate tasks, and drive continuous operational efficiency

Benefits

  • comprehensive benefits package
  • incentive and recognition programs
  • equity stock purchase
  • 401k contribution
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service