About The Position

We are seeking a Senior Staff Software Engineer (IC5 Level) to join our Data Platform Engineering organization. You will be a hands-on technical lead who will lead the design and delivery of shared, multi-tenant platform services — Postgres, queueing/streaming, and key-value stores — on Kubernetes, and serves as the technical anchor for their availability, resilience, and operating model. This position will include supporting our US Regulated Markets and requires passing a ServiceNow background screening, USFedPASS (US Federal Personnel Authorization Screening Standards), which includes a credit check, criminal/misdemeanor check and taking a drug test. Any employment is contingent upon passing the screening. Due to Federal requirements, only US citizens, US naturalized citizens or US Permanent Residents, holding a green card, will be considered.

Requirements

  • Experience leveraging or critically thinking about how to integrate AI into engineering and platform work — whether using AI-powered tooling, automating operational workflows, building agentic systems for fleet visibility and operations, or reasoning about AI’s impact on how infrastructure is built and run.
  • 12+ years of software development experience with a Bachelor's degree; OR 8+ years with a Master's degree; OR 5+ years with a PhD; OR equivalent work experience.
  • 8+ years building production software, with solid experience operating distributed systems and running stateful workloads on Kubernetes at scale.
  • A track record of leading the design and delivery of at least one shared, multi-tenant infrastructure service used broadly by other teams — a relational database service (Postgres or similar), a message queue or streaming platform (Kafka, NATS, RabbitMQ, or similar), or a key-value/cache service (Redis/Valkey, etcd, or similar) — including its HA, failover, and operating model.
  • Deep understanding of high availability and failure handling: replication topologies, leader election, quorum and consensus, split-brain avoidance, backup/restore and point-in-time recovery, and designing to explicit RPO/RTO and durability targets.
  • Strong system design skills — you can reason rigorously about consistency models, partitioning and rebalancing, capacity planning, tenant isolation, and noisy-neighbor mitigation, and communicate the trade-offs clearly to engineers and stakeholders.
  • Hands-on experience with at least one major hyperscaler (AWS, Azure, GCP), including its core compute, networking, storage, and IAM primitives.
  • Strong working knowledge of containers and Kubernetes (including stateful primitives: StatefulSets, CSI/persistent storage, PDBs, topology spread), CI/CD and GitOps-based delivery, and infrastructure-as-code.
  • Strong programming skills in Go (or strong systems-language skills with a willingness to work primarily in Go).

Nice To Haves

  • Experience building or extending Kubernetes operators that manage stateful systems (e.g., CloudNativePG, Zalando/Crunchy Postgres operators, Strimzi, Redis/Valkey operators), and opinions on when to adopt versus build.
  • Deep expertise in one of the target systems: Postgres internals (WAL, streaming/logical replication, vacuum, connection pooling with PgBouncer/PgCat, major-version upgrades); Kafka/NATS (partitioning, ISR/replication, exactly-once semantics, consumer scaling); or Redis/Valkey (cluster mode, persistence, eviction, hot-key handling).
  • Experience running data services across multiple regions or clusters — cross-region replication, failover orchestration, and DR testing.
  • Experience with zero-downtime upgrades, schema/data migrations, and fleet-wide rollouts of stateful services.
  • Experience with observability and SLOs for stateful systems — replication lag, saturation, tail latency, error budgets — and with capacity and cost management for shared infrastructure.
  • Experience designing multi-tenancy: quotas, isolation, chargeback/showback, and self-service provisioning APIs or CRDs.
  • Experience with container networking (CNI) and/or service mesh, and with workload identity, mTLS, and secrets management as applied to data services.
  • Experience with managed Kubernetes (EKS/AKS/GKE), managed data services (RDS/Aurora, Cloud SQL, MSK, ElastiCache), and infrastructure-as-code tools such as Terraform or Crossplane.

Responsibilities

  • Lead the design and delivery of shared platform services — managed Postgres, queueing/streaming, and key-value/cache — that run on Kubernetes across a large global fleet and that product teams across the company depend on, owning them from design through production operation.
  • Act as the technical lead on major initiatives within your team, breaking down ambiguous problems — “offer HA Postgres as a service to every cluster,” “make our queueing tier survive a zone loss” — into clear, executable designs with explicit availability, durability, and cost targets.
  • Define the HA, failover, disaster-recovery, and multi-region architecture for the services you own, and the Kubernetes-native automation (operators, controllers, CRDs, self-service APIs) that provisions, scales, upgrades, and fails them over without human intervention.
  • Partner with principal and distinguished engineers to align your work with the broader platform architecture and standards, and with product teams to set consumption contracts, tenancy models, and SLOs for shared services.
  • Identify technical risks early and drive them to resolution, with a strong focus on reliability, scalability, and operability — leading failure-mode analysis, game days, and post-incident reviews for stateful systems.
  • Spend significant time hands-on — designing, coding, and reviewing the core systems your team builds, such as operators, controllers, infrastructure automation, and platform services.
  • Mentor mid-level and junior engineers and raise the engineering bar through code reviews, design feedback, and pairing — particularly around distributed-systems and data-service design.

Benefits

  • health plans
  • flexible spending accounts
  • 401(k) Plan with company match
  • ESPP
  • matching donations
  • flexible time away plan
  • family leave programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service