Staff Site Reliability Engineer - US

Teleport
•$222,000 - $326,000•Remote

About The Position

Teleport is the AI Infrastructure Identity Company, focused on securing infrastructure for an AI world by giving every entity a cryptographically secured identity. The company is remote-first and globally distributed. This role is with the Teleport Cloud team, which is building its production and SaaS infrastructure from scratch. The team tackles complex security and reliability challenges to ensure customers can trust Teleport for secure and reliable access to their infrastructure. The role involves significant work in Go and requires a strong understanding of security, reliability, and engineering velocity. The company emphasizes a no-ego, transparent, and collaborative culture.

Requirements

  • Willingness to collaboratively work with Teleports’ engineers on coding challenge in Go as part of the interview process.
  • 8+ years of progressive experience in Software Engineering and/or SRE/DevOps roles.
  • Strong experience in Linux systems, networking, containers, and troubleshooting.
  • Have solid Go and production Kubernetes development experience.
  • Technical leadership experience.
  • Strong experience developing scripts, automation, submitting patches to the product codebase, or building tooling that incorporates AI agents into operational workflows.
  • Systems Observability tools: Prometheus, Grafana, Loki etc.
  • Operate and support the observability platform to maintain visibility and reliability.
  • Experience operate in a team where sound security choices are critical, and where reasoning about correctness and system invariants (e.g. formal or property-based methods) is valued.
  • Intellectual curiosity and a willingness to master new technologies.
  • Transparency, honesty, and a no-ego mindset.
  • Excellent communication skills.

Nice To Haves

  • AWS Cloud experience is preferred, GCP experience is acceptable.

Responsibilities

  • Re-engineer the core teleport product to scale globally and optimize routing latency for teams distributed around the world
  • Re-write portions of the core Teleport product to enable our goals for the cloud product
  • Build out our monitoring and observability stack to alert us to production issues and minimize false positives so we can all get a good sleep at night
  • Work on automation to tackle and eliminate the highest toil activities
  • Execute on traditional operation challenges, such as patching, scaling, backup and restore, disaster recovery, and more
  • Investigate the outages and incidents our customers experience with our product
  • Participate in the on-call rotation to ensure 24/7/365 system uptime.

Benefits

  • Extensive health coverage
  • Annual expense budget
  • Rest and recovery policies that maximize your ability to recharge
  • Investment in your future with retirement savings plans
  • Professional development opportunities
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service