Principal Distributed Systems Engineer

IonQ•Santa Clara, CA
•Onsite

About The Position

We are looking for a Principal Distributed Systems Engineer to join the team building IonQ's Network and Security Platform. As a Principal Distributed Systems Engineer, you'll be part of a cross-functional team whose mission is to lead IonQ on its journey to build the world's best quantum computers to solve the world's most complex problems — and to secure the world's networks against quantum threats. In this role, you will be responsible for setting the technical direction and end-to-end architecture of the platform's distributed backend — the ingestion pipelines, event streaming and telemetry processing, time-series and data storage layers, and APIs that give operators real-time visibility into their networks and the intelligence to make them quantum-safe. You will remain hands-on: building reference implementations and proofs of concept, reviewing critical designs and code, and turning architecture decisions into production-grade systems alongside Staff and Senior engineers. The ideal candidate is a hands-on architect and natural technical leader who has designed and scaled multi-tenant, data-intensive distributed platforms, can align technology choices with business outcomes, and is energized by making complex infrastructure simple, observable, secure, and dependable.

Requirements

  • 15+ years of software engineering experience building and operating distributed backend systems and platforms at scale, or an equivalent combination of education and experience.
  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • Demonstrated record as a hands-on architect: defining the architecture and multi-year roadmap for large-scale distributed platforms and driving them through to production.
  • Deep experience designing distributed data systems: event streaming (Kafka, Kinesis, or equivalent), high-throughput ingestion pipelines, and time-series or large-scale data stores.
  • Strong fundamentals in distributed systems: consistency models, failure modes, backpressure, idempotency, and exactly-once vs. at-least-once processing tradeoffs.
  • Hands-on proficiency in at least one production systems language (Go, Rust, Java, or Scala), with the ability to build reference implementations and review production code.
  • Extensive experience with cloud-native architecture on AWS (GCP or Azure a plus): containerization, Kubernetes, infrastructure-as-code, security, and cost optimization.
  • Experience designing multi-tenant platforms with strong data isolation, access control, and governance.
  • Experience defining observability and reliability frameworks — logging, metrics, tracing, SLOs — for mission-critical systems.
  • Established record of setting technical direction beyond a single team: leading architecture governance, influencing executive stakeholders, and mentoring senior engineers across globally distributed organizations.
  • Excellent written and verbal communication, with the ability to explain complex architecture tradeoffs to both engineering and business audiences.

Nice To Haves

  • Master's degree in Computer Science, Engineering, or a related discipline.
  • Production experience in Go and/or Rust.
  • Experience with network telemetry and network management platforms — gNMI/gRPC, NETCONF, SNMP, YANG data models — or prior work at a network equipment vendor or network management company (e.g., Juniper, Cisco, Arista, Nokia).
  • Experience with graph databases or graph data models for network topology (e.g., Neo4j, Amazon Neptune, or custom adjacency representations).
  • Experience with AI/ML platform infrastructure (MLOps / LLMOps, model governance, agentic workflows) applied to operational or network data.
  • Experience in regulated, security-conscious environments (FedRAMP, NIST, or equivalent).
  • Contributions to open-source infrastructure, streaming, observability, or networking projects, or published technical writing on distributed systems.

Responsibilities

  • Define and own the multi-year architecture and technical roadmap for the platform's distributed backend, measured by delivery against roadmap milestones and platform availability, scalability, and cost targets.
  • Lead architecture reviews and author the design standards, reference architectures, and decision records that guide how the platform — and IonQ's broader engineering organization — approaches distributed systems and data infrastructure problems.
  • Identify systemic reliability, scalability, security, or architectural gaps across the platform and drive cross-team initiatives to resolve them.
  • Partner with product, security, and engineering leadership to translate business requirements into architecture decisions, clear engineering specifications, and sequenced delivery plans.
  • Architect high-throughput ingestion and telemetry pipelines that reliably collect data from large network device fleets (e.g., gNMI/gRPC, NETCONF, SNMP streams).
  • Design event streaming architectures using Kafka or Kinesis with well-defined delivery semantics (at-least-once / exactly-once), schema enforcement and evolution, and backpressure handling.
  • Define the time-series and storage strategy (InfluxDB, TimescaleDB, or equivalent), including partitioning, retention, archival, and query performance under production load.
  • Design RESTful and gRPC APIs that expose platform data to internal consumers and external integrations, evaluated by latency SLAs, uptime, and adoption by downstream teams.
  • Build production-quality services, prototypes, and reference implementations in Go and/or Rust (or another production systems language), setting the bar for code quality and test coverage.
  • Review critical designs and code across teams to ensure correctness, performance, and operability.
  • Establish the observability architecture — structured logging, distributed tracing, and metrics (OpenTelemetry, Prometheus) — so production issues can be diagnosed and resolved within defined SLOs.
  • Define multi-tenant isolation, access control, and data-governance patterns that meet security, compliance, and data residency requirements.
  • Guide incident response for complex production issues, participate in on-call escalation, and drive systemic improvements tracked through post-mortems.
  • Set the multi-cloud deployment architecture for containerized workloads on Kubernetes, with infrastructure-as-code (Terraform) across AWS, GCP, and Azure, meeting platform availability and cost targets.
  • Mentor and coach Senior, Staff, and Senior Staff engineers; raise the bar on engineering standards, code review culture, and design practices.
  • Act as a cross-organizational technical authority on distributed systems and data infrastructure across IonQ.

Benefits

  • comprehensive medical, dental, and vision plans
  • matching 401(k)
  • unlimited PTO
  • paid holidays
  • parental/adoption leave
  • legal insurance
  • a home technology stipend
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service