Principal Site Reliability Engineer

KēSTA I.T.Culver City, CA

About The Position

An innovative technology company is seeking experienced Site Reliability Engineers to take ownership of building reliable, scalable platforms that deliver advanced 3D/4D spatial content to global users across AR/VR environments. This is a high-impact role focused on ensuring system reliability at scale, requiring deep expertise in observability, multi-tenant architectures, and data-driven operational decision-making. You will play a key role in designing and maintaining infrastructure that supports high-volume streaming workloads while meeting enterprise-grade security and compliance standards. This role partners closely with web services and platform engineering teams to implement SRE best practices, establish robust monitoring, and build infrastructure capable of supporting rapid growth and global distribution.

Requirements

  • 7+ years of experience in Site Reliability Engineering, DevOps, or related roles, with a track record of improving system reliability and operational maturity
  • Strong expertise in cloud platforms and modern infrastructure environments (e.g., AWS, containerized workloads, or similar ecosystems)
  • Experience with infrastructure automation and container orchestration (e.g., Terraform, Kubernetes or equivalent technologies)
  • Deep understanding of multi-tenant architecture, security principles, and data protection practices
  • Hands-on experience with observability tools and monitoring frameworks (e.g., Prometheus, Grafana or similar)
  • Experience implementing automated compliance and governance practices (e.g., SOC 2, GDPR, ISO 27001 or similar standards)
  • Strong leadership and mentoring capabilities, with the ability to influence engineering teams and drive adoption of reliability-focused practices

Responsibilities

  • Design, configure, and maintain cloud infrastructure using infrastructure-as-code tools (e.g., Terraform), with a focus on optimizing content delivery and CDN performance
  • Develop and execute capacity planning strategies and performance optimization initiatives for large-scale streaming platforms
  • Instrument services to monitor system health, building dashboards and alerting systems that provide actionable insights into performance and user experience
  • Define and implement observability strategies, including SLI/SLO frameworks and error budget management
  • Establish escalation protocols and participate in on-call rotations to ensure 24/7 system availability
  • Lead incident response efforts and conduct post-incident reviews to drive continuous improvement
  • Implement and promote reliability engineering practices, including deployment safety, code review standards, and operational readiness
  • Mentor engineering teams on best practices for reliability, scalability, and production operations

Benefits

  • top performance is rewarded
  • personal time is valued
  • excellence is demanded at every level
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service