Senior Production Engineer

WhatnotSan Francisco, CA
$207,000 - $290,000Hybrid

About The Position

Production Engineers are software engineers who embed with product, platform, and infrastructure teams to find and fix what breaks at scale. They are engineers who go looking for the anomaly, the bottleneck, the slow degradation nobody has noticed yet, and then own it through to a fix, working with whoever owns the system to make sure it stays fixed. Live video and real-time commerce run on the same critical path here. Every auction is live and money moves during the show, so system behavior and business outcomes are tightly coupled in a way few platforms experience. Milliseconds are visible to buyers and sellers, and the signals that matter most are often small, concentrated in a specific cohort, and invisible in aggregate. Surfacing them early, and knowing which ones are worth acting on, is the core of this work. Production Engineers embed with a team and go deep on its systems, while staying accountable for what falls between teams. Some of the most consequential problems at our scale live in the seams, in the interaction between two services that each look healthy on their own, so we expect hunting both inside your area and across it. You'll be joining early, which means broader discovery across the platform before the embedding settles, and real influence over how Production Engineering takes shape here.

Requirements

  • Bachelor’s degree in Computer Science, a related field, or equivalent work experience.
  • 6+ years building and debugging production services at scale. You may have come to that through software engineering, production engineering, SRE, or systems engineering; what matters is the depth, not the title.
  • A software engineer's identity first, paired with a pull toward the messy end of production. You want to write code and build systems, and you're the one who notices the graph that looks slightly wrong and can't let it go.
  • Strong systems and distributed systems fundamentals: failure modes, saturation, queueing behavior, cascading failure, and how Linux, networking, and storage behave under load.
  • Depth in observability and production debugging, plus experience in one or more of: capacity and performance engineering, load and resilience testing, incident command, or large-scale migrations under production constraints.
  • The ability to work well embedded. You can walk into another team's codebase, earn trust quickly, and leave the system better than you found it.
  • Fluency across languages and technologies, and comfort in cloud-native environments such as AWS or GCP with Kubernetes and infrastructure as code. Our backend is primarily Python and Elixir, with Go for performance-sensitive infrastructure, and we care more about depth and adaptability than a specific stack.

Nice To Haves

  • high-traffic, real-time, event-driven, or live streaming systems.

Responsibilities

  • Go deep with the team you embed with, raising the reliability, performance, and scalability of their systems alongside them and writing the code that gets it there.
  • Hunt anomalies across traffic, latency, error, and cost signals, both inside your area and in the seams between teams, and trace them to root cause across services, storage, and clients.
  • Connect operational signals to business impact: which degradations matter, for which cohorts, and which are safe to leave alone.
  • Work where the risk concentrates: payments, live video, search, the bidding path, and security.
  • Eliminate scale bottlenecks in production, then remove the class of problem rather than the instance.
  • Prepare for peak events through capacity modeling, load testing, and failure drills.
  • Improve observability where it's thin, so the next anomaly is found in minutes rather than quarters.
  • Put AI to work on operations: anomaly detection, on-call assistance, remediation, toil reduction. Mostly greenfield, and yours to prove out.
  • Share on-call with the team you embed with, and act as an escalation point for live production incidents.
  • Lead incident response for complex cross-team failures and drive the systemic fixes that follow.

Benefits

  • Flexible Time off Policy and Company-wide Holidays (including a spring and winter break)
  • Health Insurance options including Medical, Dental, Vision
  • Work From Home Support
  • Home office setup allowance
  • Monthly allowance for cell phone and internet
  • Care benefits
  • Monthly allowance for wellness
  • Annual allowance towards Childcare
  • Lifetime benefit for family planning, such as adoption or fertility expenses
  • Retirement; 401k offering for Traditional and Roth accounts in the US (employer match up to 4% of base salary) and Pension plans internationally
  • Monthly allowance to dogfood the app
  • Parental Leave: 16 weeks of paid parental leave + one month gradual return to work company leave allowances run concurrently with country leave requirements which take precedence.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service