Staff Infrastructure Engineer

Headway
$265,000 - $331,000

About The Position

Headway's mission is to fix the mental healthcare system by building a new system everyone can access, starting by solving the biggest barrier to care: insurance. They have automated the administrative work like credentialing, claims, and payment reconciliation. Over 75,000 providers across all 50 states use their software, serving over 1 million patients. Headway is building the best tools for therapists, reimagining the experience of finding a therapist, and investing in platform foundations to enable this at scale. They are a Series D company with over $325M in funding from investors like a16z, Accel, and Spark Capital, and are looking for exceptional people to help them achieve their mission and make this the most meaningful experience of their careers.

Requirements

  • 8 or more years in platform, infrastructure, or SRE roles at companies running significant production traffic
  • Deep AWS expertise and production ownership of compute and networking at scale (ECS, EKS, RDS, networking, IAM)
  • Strong infrastructure-as-code experience, particularly Terraform, including designing self-serve platforms for other engineering teams
  • Hands-on autoscaling and capacity engineering, and container orchestration with ECS and/or EKS
  • Track record making deploys safe and self-serve for other teams, not just your own
  • Staff-level influence: you drive decisions across team boundaries and raise the infrastructure bar org-wide without requiring management authority to do it

Nice To Haves

  • FinOps and cloud cost optimization experience
  • Kubernetes and EKS depth
  • Observability tooling at scale (Datadog)
  • Experience in healthcare or other regulated environments
  • Experience with event driven systems

Responsibilities

  • Architect and own the cloud platform that every engineer at Headway deploys on.
  • Make deploys boring, scaling automatic, infrastructure self-serve, and cost attributable.
  • Own the cloud platform Headway runs on.
  • Lead the work to isolate blast radius, make autoscaling trustworthy, and build a self-serve infrastructure platform.
  • Serve as the technical anchor for Headway's compute, networking, and deployment platform.
  • Bring Staff-level influence to an area that every engineer depends on daily.
  • Own deployment architecture and blast-radius containment, redesigning deployments to prevent mistakes in one part of the service from blocking or taking down others.
  • Continue to drive the shift toward per-service deploy isolation and functional-area slices that contain failures rather than propagating them across the platform.
  • Own the ECS and EKS footprint, evaluate broader EKS adoption for AI workloads, and design the next iteration of inter-service network connectivity.
  • Own capacity for a spiky workload: floors computed ahead of demand rather than chased by reactive scaling, self-deriving from data, with drift caught early.
  • Build the Terraform self-serve platform with guardrails so engineering teams own their standard infrastructure changes and Reliability Engineering reviews only the non-standard ones.
  • Stand up per-team cost attribution across AWS, Datadog, and LLM spend.
  • Make infrastructure costs visible and attributable so teams can make informed tradeoffs.
  • Own how the monolith behaves under load: garbage collection, event loop contention, and the runtime limits that bite first.
  • Lead the framework and package upgrades most teams defer.

Benefits

  • Meaningful work contributing to changing mental healthcare for the better
  • Opportunity to be the most meaningful experience of your career
  • Competitive salary and benefits package (implied by Series D funding and focus on exceptional people)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service