Senior Software Engineer - Observability

SnowflakeBellevue, WA
$236,000 - $339,200Hybrid

About The Position

Snowflake's Data Cloud processes exabytes of data across multi-cloud global environments every day. Delivering seamless reliability and real-time visibility to thousands of global enterprise customers requires an Observability Platform built on hyper-scalable backend distributed systems. We are seeking a Senior Software Engineer to own key components of our AI native External Observability Platform. In this role, you will contribute to the technical road map for customer-facing telemetry, system metrics, audit logs, distributed tracing, and actionable operational insights. You will build high-throughput, low-latency infrastructure capable of ingesting, processing, and serving petabytes of telemetry data with strict SLA guarantees. You will join a team of world-class engineers in our Bellevue, WA office. To be successful, you must be deeply technical, capable of leading complex technical projects, and skilled at collaborating with the brightest technical minds in the industry.

Requirements

  • 7-12 years of professional experience building infrastructure and backend distributed systems at scale using languages such as Go, C++, Java, or Rust.
  • Proven track record of developing, deploying, and maintaining hyper-scale distributed systems in public cloud environments (AWS, Azure, or GCP).
  • Deep theoretical and practical CS fundamentals (data structures, algorithms, concurrency patterns, storage engines, distributed consensus, networking protocols).
  • Strong experience with infrastructure-as-code (IaC) tools such as Terraform, or Pulumi.
  • Hands-on expertise with time-series databases, high-cardinality metric stores, distributed tracing frameworks (OpenTelemetry, Jaeger), or log streaming systems (Kafka, Flink, ClickHouse, Prometheus/Thanos).
  • Superior communication and collaboration skills.

Nice To Haves

  • Massive Scale Experience: Prior experience in high-performance computing (HPC) or handling global installations processing petabytes of telemetry per day.
  • Customer-Facing Observability: Experience building external or customer-exposed observability tools, APIs, and analytics dashboards with strict latency and access-control constraints.
  • Network & Systems Deep Dive: Deep operational understanding of high-performance Linux kernel tuning, ebpf, eBPF-based profiling, or low-level network optimization.
  • Multi-Cloud Expertise: Prior experience running cloud-agnostic platform infrastructure across AWS, Azure, and GCP simultaneously.

Responsibilities

  • Develop and Scale Distributed Infrastructure: Design and implement key components of Snowflake’s External Observability platform capable of handling trillions of events per day across multi-cloud deployments (AWS, Azure, GCP).
  • Build High-Throughput Engines: Write scalable, reliable, and testable backend services to process time-series data, high-cardinality metrics, and distributed traces at massive scale.
  • Automate Infrastructure Lifecycle: Practice infrastructure-as-code (IaC) using tools like Terraform to deliver self-healing, automated telemetry pipelines and dynamic monitoring topology.
  • Drive Platform Adoption & Diplomacy: Collaborate closely with Application Engineering, Security, and Customer Support teams to make systems measurable, translate telemetry into actionable customer-facing insights, and resolve cross-organizational technical dependencies.
  • Incident Escalation & Root Cause Analysis: Serve as an expert troubleshooter for critical, complex system failure modes across large distributed clusters, conducting deep-dive post-mortems and building automation to eliminate repeat incidents.

Benefits

  • Confidentiality and security standards for handling sensitive data
  • Abide by the company’s data security plan
  • Keep customer information secure and confidential
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service