Staff+ Software Engineer - Backend

ResolveAISan Francisco, CA
Onsite

About The Position

Resolve AI is solving the problem of software maintenance and production troubleshooting, which significantly impacts engineering velocity. We are building a transformative, autonomous AI Production Engineer designed to investigate and fix complex system issues end-to-end. Our founders, who were instrumental in creating OpenTelemetry and led Splunk Observability, have a strong track record with successful exits to Splunk and VMware. We have secured over $190M in funding from prominent investors and notable figures in the tech industry. Joining Resolve AI at this stage presents a unique opportunity to work at an AI-native company that is redefining the future of engineering work.

Requirements

  • Ideally with 1+ years operating at a Staff or equivalent technical leadership level.
  • Strong backend engineering experience in building and deploying scalable, high-performance services in high-concurrency production environments.
  • Experience with cloud platforms (especially AWS) and Kubernetes.
  • Solid understanding of databases and messaging systems for cloud-native applications.
  • Bias for action: ability to quickly get initial versions to users and iterate based on feedback.
  • Willingness to delve into technical details, debug complex problems, analyze data, and get hands-on.
  • Proficiency in using AI tools to enhance productivity and eagerness to explore LLM-powered systems.

Nice To Haves

  • Prior SRE, on-call, or incident response experience.
  • Familiarity with observability tools and the internals of developer-facing tools.

Responsibilities

  • Build simulation environments for training and evaluating AI agents on realistic SRE tasks, including debugging, incident response, infrastructure changes, and on-call duties.
  • Collaborate with Research Scientists and domain experts to translate real-world SRE scenarios into reproducible and challenging graded environments.
  • Take ownership of the architecture and implementation of major subsystems, from initial design through ongoing evolution, including infrastructure, tooling, data pipelines, and scoring.
  • Design, build, and scale backend services and APIs for generating, running, and scoring scenarios in high-concurrency environments.
  • Ensure reliability, scalability, and performance are integrated into every system layer, with a particular focus on distributed systems.
  • Iteratively develop and refine systems by shipping early versions, monitoring performance, identifying issues, and making improvements.
  • Focus on the realism of environments, ensuring they accurately reflect production failures rather than sanitized versions.
  • Leverage AI tools across the entire stack, including code generation, scenario generation, and debugging, to enhance productivity.

Benefits

  • Competitive Pay Packages
  • Comprehensive Medical, Dental, and Vision Insurance
  • Monthly Housing Stipend
  • Flexible (Unlimited) Paid Time Off
  • Visa Sponsorship & Immigration Support
  • 401(k) Plan
  • Parental Leave
  • Discretionary Tech Benefit Stipend
  • Daily in-office Lunches and Dinners
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service