Staff Site Reliability Engineer

BlueskyFully Remote - Overlap with PST working hours ,
$200,000 - $270,000Remote

About The Position

Bluesky's mission is to transition the social web from platforms to protocols. We're building a federated social network called AT Protocol where users have more power. We're looking for a Staff Site Reliability Engineer to help design, implement, and operate the infrastructure that powers Bluesky and atproto. This is a hands-on role for someone who has operated high-scale production systems, understands how distributed systems fail, and wants to build the operational foundation for an open social network.

Requirements

  • 10+ years experience operating high-scale production systems, including bare metal
  • Strong fundamentals in Linux, networking, storage, databases, and distributed systems
  • Built and operated high-scale systems where correctness, latency, throughput, and availability were critical
  • Can write production-quality software in Go
  • Comfortable debugging across application code, operating systems, databases, networks, and hardware
  • Experience with observability systems, alert design, incident response, capacity planning, kubernetes, and production automation
  • Like working on very small, fast-moving teams at a startup
  • Have read the AT Protocol docs, feel aligned with the mission, and want to contribute!

Responsibilities

  • Work across bare-metal systems, cloud services, data infrastructure, observability, incident response, capacity planning, and reliability engineering for systems serving millions of users.
  • Own reliability, availability, and operational excellence for our production systems, including observability, incident response, deployment, and rollback systems.
  • Improve production readiness for services, migrations, and infrastructure changes.
  • Develop software that pushes the state of the art in performance, automation, observability, and other areas.
  • Scale systems running on dense, latest-generation, bare-metal servers in our own colocation facilities.
  • Reduce toil through automation, tooling, and thoughtful engineering practices.
  • Partner with engineers across all our teams to help design services with strong operational characteristics.
  • Lead incident reviews and turn contributing factors into concrete engineering improvements as we practice continuous improvement.
  • Perform capacity planning and cost management across compute, storage, database, and networking workloads.
  • Manage various vendor relationships to ensure we can provide high quality services at a reasonable TCO.
  • Mentor engineers on reliability, operability, debugging, and distributed systems practices and help define a culture of operational excellence across the org.

Benefits

  • health insurance
  • dental insurance
  • vision insurance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service