Senior Site Reliability Engineer

GridCARERedwood City, CA
$180,000 - $230,000Hybrid

About The Position

We're looking for a Senior SRE to own the reliability, scalability, and observability of our production systems. You'll work closely with platform and data engineering to keep high-throughput, data-intensive services running at the availability our customers (utilities, data center operators) require.

Requirements

  • 5+ years in SRE, DevOps, or infrastructure engineering roles
  • Deep experience with Kubernetes, Terraform/IaC, and cloud platforms (AWS Preferred)
  • Strong scripting/programming ability (Python, Bash)
  • Observability Experience (Prometheus, Grafana, Datadog)
  • Track record of running on-call for production systems and leading incident response
  • Experience with CI/CD pipelines (Github Actions) and infrastructure automation
  • Solid understanding of networking, distributed systems, and database reliability
  • Comfortable operating in a fast-moving startup environment with ambiguity

Nice To Haves

  • Experience with data-intensive or real-time processing systems
  • Background in energy, climate tech, or critical infrastructure
  • Experience scaling infrastructure through hypergrowth

Responsibilities

  • Design and operate infrastructure on AWS using Terraform and Kubernetes
  • Build monitoring, alerting, and observability (Prometheus, Grafana, Datadog, or similar) with meaningful SLOs/SLIs
  • Automate away toil — deployment pipelines, capacity management, self-healing systems
  • Partner with engineering on architecture reviews to catch reliability and scalability risks before they ship
  • Manage database and data pipeline reliability for large-scale, real-time grid data processing
  • Drive security and compliance best practices across infrastructure

Benefits

  • Competitive salary
  • performance bonus
  • equity
  • Comprehensive health, dental, and vision coverage
  • Lunch provided three days a week in office
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service