Lead SWE, AI Dev/IT Operations - Remote or Hybrid in DC or MN

UnitedHealth GroupMinnetonka, MN
$112,700 - $193,200Hybrid

About The Position

Optum Tech is a global leader in health care innovation. Our teams develop cutting-edge solutions that help people live healthier lives and help make the health system work better for everyone. From advanced data analytics and AI to cybersecurity, we use innovative approaches to solve some of health care’s most complex challenges. Your contributions here have the potential to change lives. Ready to build the next breakthrough? Join us to start Caring. Connecting. Growing together. We're seeking a skilled Lead Software Engineer to help build intelligent automation, predictive analytics, and agentic AI capabilities for IT Operations. In this role, you will design, build, and help operationalize machine learning models and AI-driven workflows that improve observability, speed up incident response, and support self-healing infrastructure across the enterprise. You'll work closely with senior engineers, architects, and data scientists to bring these capabilities into production. You’ll enjoy the flexibility to work remotely from anywhere within the U.S. as you take on some tough challenges. For all hires in the Minneapolis or Washington, D.C. area, you will be required to work in the office a minimum of four days per week.

Requirements

  • Bachelor's degree in Computer Science, Data Science, Engineering, or a related field, or equivalent practical experience
  • 3+ years of experience building and deploying AI/ML solutions in production environments
  • 3+ years of experience in Python and SQL, with hands-on experience using distributed data processing frameworks such as Spark
  • 2+ years of experience applying ML to IT operations, infrastructure, reliability engineering, or observability domains
  • 2+ years of experience with cloud platforms (Azure, AWS, or GCP), including hybrid and cloud native architectures
  • 2+ years of experience with MLOps practices, CI/CD pipelines, model monitoring, and lifecycle management
  • 1+ years of experience with NLP, Large Language Models (LLMs), Generative AI, and modern AI frameworks
  • 1+ years of experience with containerization and orchestration technologies such as Docker and Kubernetes

Nice To Haves

  • Hands-on experience contributing to agentic or autonomous AI systems in operational contexts
  • Experience in healthcare, regulated industries, or large-scale enterprise platforms
  • Experience working in agile, product-driven environments delivering scalable, high-impact solutions
  • Working knowledge of IT operations concepts, including incident management, CMDB, monitoring, alerting, and reliability engineering
  • Familiarity with Responsible AI and governance frameworks, including fairness, transparency, and explainability
  • Solid technical communication skills, including the ability to explain AI concepts to non-technical stakeholders

Responsibilities

  • Design, build, and maintain machine learning models and AI-driven automation for IT operations use cases, including predictive analytics, intelligent automation, and support for autonomous remediation
  • Build and maintain data pipelines that ingest structured and unstructured data from observability and ITSM platforms (e.g., Splunk, Dynatrace, ServiceNow, logs, metrics, traces)
  • Develop and maintain feature sets that support anomaly detection, predictive alerting, capacity forecasting, and operational intelligence use cases
  • Partner with data scientists and platform teams to deploy and productionize ML models, and help ensure they perform reliably at scale
  • Integrate ML and AI outputs into orchestration and automation platforms (e.g., Jenkins, Interlink) to support closed-loop, self-healing operational workflows
  • Build components of agentic AI solutions that reason, plan, and execute actions across IT operations, including incident triage, root cause analysis, and resolution
  • Follow established MLOps standards, model lifecycle practices, and data governance requirements aligned with enterprise security, privacy, and compliance policies
  • Use cloud native and hybrid architectures to support real-time analytics, scalable model inference, and resilient AI platforms
  • Work with engineering, operations, security, and business stakeholders to understand requirements, document your work, and contribute reusable patterns for AI-driven operations
  • Apply ethical AI principles, including fairness, transparency, explainability, and accountability, throughout the development lifecycle

Benefits

  • Comprehensive benefits package
  • Incentive and recognition programs
  • Equity stock purchase
  • 401k contribution
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service