Member of Technical Staff - Reliability Engineering

FireworksSan Mateo, CA
$240,000 - $290,000

About The Position

Fireworks AI is a leader in inference and training for open models, aiming to make these models fast, cheap, and dependable. This role focuses on ensuring the platform runs reliably as it scales. The Reliability Engineering team collaborates across cloud infrastructure, AI systems, and product teams to ensure seamless integration, graceful failure handling, and robust performance under load.

Requirements

  • Systems fundamentals: 5+ years with Linux internals, system performance troubleshooting, and networking fundamentals (TCP/IP, HTTP, gRPC).
  • Software engineering: 5+ years in Python, Go, C++, or Rust, writing production-grade tools and systems code.
  • Cloud-native operations: Operating and debugging Kubernetes, Terraform, and Docker in high-throughput production.
  • Distributed systems: Experience with high-throughput control planes, microservices, or multi-region setups.
  • Reliability fundamentals: Understanding of fault-tolerant design, SLO/SLA management, automated failover, and high-availability architecture.
  • Influence without authority: Ability to drive adoption of standards through credibility and useful tooling.
  • Breadth over comfort: Willingness to investigate unfamiliar parts of the stack when problems cross system boundaries.
  • Education: Bachelor's or Master's in Computer Science, Computer Engineering, or equivalent practical experience.

Nice To Haves

  • Observability tooling: Experience with Prometheus, Grafana, OpenTelemetry, and effective alerting systems.
  • GPU and ML infrastructure exposure: Familiarity with GPUs, inference serving, or distributed training.
  • AI-assisted operations: Experience building agents or LLM-based tooling for investigation, triage, or automation.
  • Open source background: Contributions to infrastructure, systems, or ML serving projects.
  • Startup agility: Comfort working in an environment where pragmatism and teamwork are prioritized over rigid processes.

Responsibilities

  • Define reliability standards, including SLOs, error budgets, and production readiness criteria, informed by system instrumentation and telemetry.
  • Own the reliability toolchain, encompassing logging and telemetry pipelines, alerting standards, failure injection, load testing, self-healing automation, and AI-assisted investigation tooling.
  • Ensure customer experience is maintained by addressing per-service reliability and ensuring failures have owners and fixes, even when individual systems are within their SLOs.
  • Identify and resolve failures that occur at the seams between systems, such as problematic retries, non-composing timeouts, or unmapped dependencies.
  • Manage incident response, including coordinating live production issues, conducting blameless postmortems, and tracking follow-up actions to completion.
  • Reduce operational toil through automation to prevent on-call load from becoming unsustainable with growth.
  • Partner with various teams, including cloud infrastructure (capacity, multi-region risk), inference and training (serving/training stack failures), performance (zero-downtime rollouts), and product/control plane (customer-facing reliability).

Benefits

  • Solve Hard Problems: Tackle challenges at the forefront of AI infrastructure.
  • Build What’s Next: Work with bleeding-edge technology.
  • Ownership & Impact: Join a fast-growing team where your work directly shapes the future of AI.
  • Learn from the Best: Collaborate with world-class engineers and AI researchers.
  • Equal-opportunity employer status.
  • Inclusive environment.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service