Staff Site Reliability Engineer

Hippocratic AIMenlo Park, CA

About The Position

We're looking for a Senior Site Reliability Engineer who is equally at home writing production software and running the infrastructure it lives on — and who wants to take ownership of one of the hardest, highest-leverage problems on our platform: intelligently managing a large fleet of GPU-backed models. We run nearly 30 models across heterogeneous hardware, and keeping that fleet fast, reliable, and cost-effective is a serious engineering challenge. You'll build the GPU management and scheduling platform that sits at the center of it — collecting utilization and load metrics, interpreting what they actually mean, and using them to make real-time decisions about admission control and scaling. The goal: route and schedule inference calls so we use our capacity efficiently without exceeding it, and scale model replicas up and down automatically as demand shifts. This is a senior role for someone with a decade in the field who can move fluidly between systems engineering and software development, and who is excited to own a complex, evolving system end to end.

Requirements

  • 10+ years of professional experience across site reliability / DevOps engineering and software engineering
  • Computer Science Degree Required from a top CS program.
  • Strong software engineering fundamentals — you build orchestration and scheduling systems in Python and/or Go, not just configure off-the-shelf tools
  • Experience designing systems that make decisions from operational metrics — collecting signals, interpreting them, and driving control loops such as autoscaling, load shedding, or admission control
  • Deep experience with infrastructure automation and CI/CD (Terraform, GitLab CI/CD, or similar)
  • Hands-on production experience with at least one major cloud platform (AWS, GCP, or Azure)
  • Strong knowledge of containerization and orchestration (Docker, Kubernetes)
  • Experience with monitoring and logging stacks (ELK, Grafana, Datadog, or similar)
  • Familiarity with secrets management and security tooling (HashiCorp Vault, AWS KMS, Azure Key Vault)
  • Excellent problem-solving skills and the ability to work both independently and collaboratively
  • Strong communication and interpersonal skills

Nice To Haves

  • Experience managing GPU fleets or scheduling workloads across heterogeneous accelerators
  • Familiarity with ML inference serving and model deployment (e.g. Triton, KServe, Ray Serve, or similar)
  • Experience with Kubernetes autoscaling internals (HPA/VPA, custom metrics, custom controllers)
  • Experience implementing HIPAA and SOC 2 compliance
  • Experience operating in an HPC environment
  • Bachelor's or Master's in Computer Science, Computer Engineering, or a related field

Responsibilities

  • Design and build our GPU management and scheduling platform — the system that decides when, where, and how inference calls run across a fleet of ~30 models on heterogeneous hardware
  • Build the metrics pipeline that collects GPU load and utilization data, and the logic that turns those signals into decisions
  • Implement admission control to protect capacity — deciding when to accept, queue, or shed inference requests so we operate within fleet limits
  • Build autoscaling that adjusts the number of model replicas in response to real-time demand and utilization
  • Develop cloud orchestration systems and operators in Python and Go to manage the model fleet
  • Architect and operate scalable, fault-tolerant, secure production systems on AWS, GCP, or Azure
  • Design and build infrastructure automation and deployment pipelines (Terraform, CI/CD) as first-class software
  • Stand up and maintain monitoring, logging, and alerting that keep the platform reliable and performant
  • Develop and enforce security and compliance policies appropriate to a healthcare AI platform
  • Partner with engineers and research scientists to diagnose and resolve complex infrastructure, deployment, and operational issues
  • Mentor engineers and raise the technical bar across the team

Benefits

  • Reinvent healthcare with AI that puts safety first.
  • Work with the people shaping the future.
  • Backed by the world’s leading healthcare and AI investors.
  • Build alongside the best in healthcare and AI.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service