Manager, Site Reliability Engineer

Twist Bioscience•South San Francisco, CA
•$210,383 - $243,210•Hybrid

About The Position

We are seeking an experienced and hands-on Site Reliability Engineering (SRE) Manager to lead our Site Operations and infrastructure initiatives. This role is responsible for ensuring the reliability, scalability, security, and performance of our cloud infrastructure and critical applications. The ideal candidate combines strong technical expertise with leadership experience and a passion for operational excellence. You will lead efforts across cloud infrastructure, observability, automation, networking, and platform reliability while partnering closely with engineering and product teams to support business-critical applications.

Requirements

  • 5–7+ years of experience in Site Operations, Site Reliability Engineering, Infrastructure Engineering, or related operational roles
  • Prior experience leading or mentoring engineering or operations teams
  • Strong hands-on experience with cloud infrastructure platforms (AWS, Azure, and/or GCP)
  • Experience managing Kubernetes (K8s) environments in production
  • Strong understanding of networking concepts including DNS, load balancing, firewalls, routing, and CDN technologies
  • Experience with Akamai CDN administration and optimization
  • Experience with observability and monitoring platforms, including Splunk
  • Experience with storage infrastructure and cloud-native storage solutions
  • Strong scripting and automation experience (Python, Bash, Terraform, or similar)
  • Experience implementing and maintaining CI/CD pipelines
  • Strong troubleshooting and incident management skills in complex distributed systems
  • Excellent communication and cross-functional collaboration skills

Nice To Haves

  • Experience managing hybrid or multi-cloud environments
  • Experience with infrastructure-as-code and GitOps methodologies
  • Familiarity with security best practices and compliance frameworks
  • Experience supporting highly available enterprise or SaaS platforms
  • Strong understanding of container security and cloud-native architectures

Responsibilities

  • Lead and manage the Site Operations / SRE function
  • Own cloud infrastructure architecture, operations, scalability, and optimization
  • Ensure high availability and reliability of production applications and services
  • Drive operational excellence through automation, monitoring, and incident management
  • Develop and maintain observability platforms including logging, metrics, alerting, and tracing
  • Manage Kubernetes-based infrastructure and container orchestration environments
  • Oversee networking, DNS, CDN, and storage infrastructure across cloud environments
  • Collaborate with software engineering teams to improve deployment processes, system resilience, and performance
  • Establish and improve CI/CD pipelines and infrastructure-as-code practices
  • Lead incident response, root cause analysis, and post-incident remediation efforts
  • Manage vendor and platform relationships where applicable
  • Mentor and develop SRE / Site Operations engineers and foster a culture of reliability and continuous improvement
  • Define and implement infrastructure security and compliance best practices

Benefits

  • bonus
  • equity
  • generous benefits package
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service