Site Reliability Engineer

FabricNew York, NY

About The Position

We are looking for a Site Reliability Engineer to help us evolve and safeguard the infrastructure powering healthcare experiences for millions of patients, all while keeping operational toil to a minimum for the Fabric tech community.

Requirements

  • 5+ years of experience in SRE or Platform Engineering roles managing production environments at scale.
  • Expert technical depth in AWS (EKS, EC2, RDS, S3) and production-grade Kubernetes management.
  • Proficiency with modern tooling including Terraform (IaC), Datadog (Observability), Helm (Release Management), and GitHub Actions (CI/CD).
  • Solid coding and scripting skills in Python, Bash, or Go.
  • A "rigor-first" mindset with a dedication to HIPAA-compliant, high-availability architecture.

Nice To Haves

  • Preferred experience building agentic workflows or AI-assisted tooling to drive operational efficiency.

Responsibilities

  • Designing, deploying, and maintaining Kubernetes (EKS) clusters for enterprise-grade availability.
  • Optimizing the footprint of infrastructure objects across core AWS services (EC2, RDS, S3) for performance, cost, and reliability.
  • Evolving a scalable infrastructure management platform, with the right interfaces and guardrails to maximize engineering agency at minimal cognitive load.
  • Defining golden paths for workload orchestration and integration with infrastructure dependencies, ensuring the right way to do things is also the easiest.
  • Providing robust, reusable GitHub Actions components to streamline the software delivery lifecycle.
  • Developing internal tools that replace manual operations with intelligent, autonomous systems.
  • Exploring and deploying agentic workflows for AI-assisted runbooks that automate complex, error-prone procedures and repetitive tasks.
  • Driving the evolution of observability practices by maintaining reliable mechanisms to collect the metrics, traces, and logs needed to meet SLOs.
  • Leading incident response efforts and facilitating the blameless postmortems that help systematically reduce recovery time (MTTR).
  • Defining and monitoring platform-level SLIs and SLOs to ensure it consistently meets rigorous healthcare performance standards.
  • Ensuring every piece of infrastructure is continuously compliant with HIPAA and other critical healthcare regulatory requirements.
  • Reviewing architectural decisions and technical proposals, asking incisive questions to surface technical and organizational risks before they become incidents.
  • Mentoring engineers across the company on reliability best practices and contributing a clinical-safety perspective to cross-functional design reviews.

Benefits

  • medical
  • dental
  • vision
  • unlimited PTO
  • 401(k) plan
  • stock options
  • bonuses
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service