Principal Production Engineer, Database Infrastructure

GitHub, Inc.UNAVAILABLE, UNAVAILABLE
Remote

About The Position

GitHub is looking for a Principal Production Engineer to help scale our data platform to millions of developers. We are software engineers who specialize in reliability, working with technical partners, leading design reviews, writing SDKs and tooling to build against, and shaping how our data platform is used to prevent scaling problems before they reach production. The team is highly distributed across geographies and time zones, and you will thrive in an environment of remote work and asynchronous communication.

Requirements

  • 11+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python OR Associate's Degree in Computer Science, Electrical Engineering, Electronics Engineering, Math, Physics, Computer Engineering, Computer Science, or related field AND 10+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python OR Bachelor's Degree in Computer Science or related field AND 9+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python OR Master's Degree in Computer Science, Electrical Engineering, Electronics Engineering, Math, Physics, Computer Engineering, Computer Science, or related field AND 7+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python. OR Doctorate in Computer Science, Electrical Engineering, Electronics Engineering, Math, Physics, Computer Engineering, Computer Science, or related field AND 5+ years experience in Software Engineering, Computer Science, or related technical discipline with proven experience maintaining and delivering production software coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Go, Ruby, Rust, or Python. OR equivalent experience.
  • 5+ years experience operating large-scale distributed systems in production, including participation in an on-call rotation.

Nice To Haves

  • Excitement about building, operating, and maintaining resilient, scalable systems that impact a global community of users with the ability to break down complex systems into manageable components.
  • A track record of partnering with product and feature teams and materially changing the reliability and scalability of what they ship.
  • Ability to influence engineering decisions and proactively engage in system design conversations.
  • Experience running stateful services on managed cloud data stores, specifically Azure Cosmos DB and Azure SQL Database, or equivalents such as DynamoDB, Aurora, or Cloud Spanner with a focus on partition and index design, consistency and isolation tradeoffs, throughput provisioning, and hot-partition diagnosis.
  • Experience diagnosing and resolving application-level scalability problems: N+1 query patterns, hot partitions, unbounded fan-out, cache stampedes, and data access patterns that don’t scale.
  • A track record of building internal platforms, SDKs, or developer tools adopted across an engineering organization, written in production-grade Go, Python, Ruby, or Rust
  • Deep familiarity with the failure modes of large-scale systems, both in the application and in the platform beneath it e.g. cascading failures, retry storms, thundering herds, partial outages, throttling and quota limits, control plane outages, noisy neighbors, and the patterns that mitigate them.
  • Experience leading large-scale cloud migrations of live, high-traffic services e.g. dual-write and backfill strategies, traffic shifting, correctness verification, and rollback under load.
  • Effective communication skills and willingness to pair on problems, brainstorm in public, and enthusiastically engage with your teammates in group problem solving.

Responsibilities

  • Partner with product and feature teams by leading design reviews and shaping how they model, access, and scale their data so the applications they build are performant, available, and operable at GitHub's scale
  • Design and ship SDKs, client libraries, and the tooling applications are built on
  • Set reliability strategy for GitHub's database platforms across multiple systems, defining SLOs and operational standards
  • Write technical documentation and advocate for the health and quality of the systems the team builds
  • Participate in an on-call rotation and respond to incidents as needed
  • Develop and design plans for disaster recovery, load shedding, and regional failover

Benefits

  • competitive pay
  • generous learning and growth opportunities
  • excellent benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service