Staff Software Engineer, Observability

RobinhoodMenlo Park, CA
Hybrid

About The Position

Robinhood is building an elite team to apply frontier technologies to the world's biggest financial problems. They are looking for bold thinkers and sharp problem-solvers who are wired to make an impact. Robinhood is a place where ambitious people do the best work of their careers, operating as a high-performing, fast-moving team with ethics at the center of everything they do. The Observability team's mission is to build and own Robinhood's full-stack observability platform, which is the foundation that keeps every product, service, and customer experience running reliably at scale. They design and operate systems that give engineers deep visibility into how Robinhood's infrastructure behaves, ensuring that when something goes wrong, the right people know immediately and can act fast. Their work spans metrics, logs, distributed tracing, and alerting pipelines, and they partner closely with the Robinhood Command Center to ensure their observability systems meet or exceed 99.9% uptime. The observability infrastructure is considered a product, not just a tool. As a Staff Software Engineer on the Observability team, you will be the technical lead shaping the roadmap for how Robinhood observes itself at scale. You will own the observability control plane end-to-end, making architectural decisions that directly impact the reliability and operational health of Robinhood's entire product surface. You'll lead a team of six engineers, drive the strategy for cost-efficient telemetry ingestion, build in-house solutions where off-the-shelf tooling falls short, and integrate AI-driven approaches to accelerate progress toward full-stack observability. This is a high-ownership, high-visibility role where your technical decisions set the direction for the entire organization! This role is based in our Menlo Park, CA office, with in-person attendance expected at least 3 days per week. Robinhood believes in the power of in-person work to accelerate progress, spark innovation, and strengthen community. Their office experience is intentional, energizing, and designed to fully support high-performing teams.

Requirements

  • 8+ years of software engineering experience, with a proven track record of owning and delivering large-scale observability or infrastructure platform initiatives.
  • Deep expertise in Kubernetes and public cloud environments (AWS preferred), with the ability to architect and operate distributed systems at scale regardless of specific vendor tooling.
  • Strong coding proficiency in one or more languages (Go, Python, or similar) with experience integrating observability agents, libraries, and instrumentation directly into production codebases.
  • Demonstrated experience owning an observability control plane or telemetry pipeline — including ingestion cost management, cardinality control, and signal routing — in a high-traffic production environment.

Nice To Haves

  • Experience with the Vector data pipeline (or equivalent high-throughput log/metrics routing tools) and familiarity with tools such as Prometheus, Grafana, Honeycomb, Humio, or Sentry is a plus.

Responsibilities

  • Define and execute the full-stack observability roadmap, establishing a clear technical vision for metrics, logs, traces, and alerting infrastructure across Robinhood's engineering organization.
  • Own and evolve the observability control plane, including telemetry ingestion pipelines, cost management strategies, and the tooling that ensures observability components remain highly available.
  • Lead and collaborate with a team of six engineers to build scalable, self-service observability solutions that enable product and infrastructure teams to move faster with greater confidence.
  • Establish and maintain SLOs for observability systems that meet or exceed Robinhood's 99.9% uptime target, ensuring the observability platform is as reliable as the services it monitors.
  • Partner with the Robinhood Command Center and engineering teams across the organization to align on dependency mapping, incident response workflows, and observability standards.

Benefits

  • Challenging, high-impact work to grow your career
  • Performance driven compensation with multipliers for outsized impact, bonus programs, equity ownership, and 401(k) matching
  • Top Tier benefits to fuel your work, including 100% paid health insurance for employees with 90% coverage for dependents
  • Access to the best AI tools on the market and continuous AI skill-building for every employee, technical or not
  • Lifestyle wallet - a highly flexible benefits spending account for wellness, learning, and more
  • Employer-paid life & disability insurance, fertility benefits, and mental health benefits
  • Time off to recharge including company holidays, paid time off, sick time, parental leave, and more!
  • Exceptional office experience with catered meals, events, and comfortable workspaces.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service