Senior Platform Engineer

HealiomGrand Rapids, MI

About The Position

The role You're building the platform that makes AI agents trustworthy enough to operate in healthcare. Why this matters When 30+ AI agents operate autonomously — coordinating schedules, handling credentialing, working alongside clinicians — infrastructure is what makes the difference between "colleague" and "liability." Observability isn't just for debugging code. It's for understanding agent behavior. Knowing when Holmes is drifting. Catching failures before patients do. What you'll build CI/CD that ships agent capabilities safely, not just code Observability that tracks agent behavior, not just service health Infrastructure that scales with agent count Automation so humans aren't the bottleneck Runbooks and incident response that let us sleep This is infrastructure for agents to operate on. That changes what platform engineering means.

Requirements

  • 7+ years DevOps/SRE/Platform Engineering, cloud-native
  • Strong Python/Bash
  • Kubernetes, Docker, Terraform
  • Grafana, Prometheus, Jaeger, OpenTelemetry
  • AWS depth
  • Experience being paged, fixing issues, and implementing preventative measures.

Nice To Haves

  • Healthcare, HIPAA, or SOC2 experience
  • AI/ML infrastructure (model serving, GPU clusters, vector DBs)
  • Opinions about what "agent observability" should mean
  • Experience as an early employee (e.g., employee #1-10) at a company.

Responsibilities

  • Build the platform that makes AI agents trustworthy enough to operate in healthcare.
  • Build CI/CD that ships agent capabilities safely, not just code.
  • Build observability that tracks agent behavior, not just service health.
  • Build infrastructure that scales with agent count.
  • Implement automation so humans aren't the bottleneck.
  • Develop runbooks and incident response that let us sleep.

Benefits

  • Real equity
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service