ML Ops Engineer

Eli Lilly and CompanySan Francisco, CA
Hybrid

About The Position

About the Lilly and NVIDIA Partnership Lilly and NVIDIA are launching a new AI co-innovation lab in the heart of Silicon Valley — an up-to-$1 billion, multi-year commitment to solve drug discovery’s toughest challenges. The lab brings Lilly scientists, technologists, chemists and biologists together with NVIDIA engineers under one roof. Together, we are building purpose-built foundation and frontier AI models trained on Lilly data at scale, tightening the feedback loop between automated wet labs and computational dry labs, designing the next generation of medicines for millions of patients across the globe. What You’ll Be Doing As an ML Ops Engineer, you build and operate the platforms that run the end-to-end machine learning lifecycle. You enable reliable model deployment, operation, monitoring, retraining, and reproducibility at scale. You optimize infrastructure and GPU resources to support research and discovery workloads. You will work closely with engineering and scientific teams to deliver production-ready AI capabilities.

Requirements

  • Strong Python skills and experience working with machine learning frameworks such as PyTorch, JAX, or TensorFlow.
  • Experience deploying, operating, and scaling production machine learning platforms, including model serving, monitoring, and large-scale inference workloads.
  • Experience with MLOps platforms and tools like MLflow, Weights & Biases, KServe, or similar technologies.
  • Proficiency with containerization, orchestration, and distributed compute environments (Docker, Kubernetes, Slurm, Ray).
  • Experience operating large-scale AI platforms that deploy, host, and optimize machine learning models for production use, using technologies such as Triton, vLLM, or TensorRT-LLM.
  • Experience with infrastructure automation and CI/CD practices using tools such as Terraform, Ansible, GitHub Actions, or related.
  • Experience supporting cloud platforms (AWS, Azure, or GCP) and on-premises GPU infrastructure.
  • Knowledge of observability and operational monitoring, including metrics, logging, tracing, and performance tuning.
  • Ability to identify and address system, infrastructure, and model performance issues through automation and continuous improvement.
  • Ability to collaborate effectively with research scientists, AI engineers, and infrastructure teams in a fast-paced environment.
  • Bachelor’s in Computer Science, Engineering, Statistics, Mathematics, or a related technical field
  • 4+ years of experience in machine learning engineering, ML Ops, or platform engineering.

Responsibilities

  • Lead the operational lifecycle of ML models, including deployment, monitoring, and ongoing reliability.
  • Operate and optimize large-scale inference platforms that support scientific discovery and AI workloads.
  • Ensure models can be deployed, scaled, monitored, and maintained in production environments.
  • Test, refine, and improve model accuracy.
  • Work with data scientists, business analysts and partners to integrate ML models into broader strategies.
  • Automate the platform with infrastructure-as-code and CI/CD, and document it well enough that someone else can operate it.

Benefits

  • company bonus (depending, in part, on company and individual performance)
  • company-sponsored 401(k)
  • pension
  • vacation benefits
  • medical, dental, vision and prescription drug benefits
  • flexible benefits (e.g., healthcare and/or dependent day care flexible spending accounts)
  • life insurance and death benefits
  • certain time off and leave of absence benefits
  • well-being benefits (e.g., employee assistance program, fitness benefits, and employee clubs and activities)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service