Staff Backend Software Engineer, Developer Productivity Engineering

TapestryMountain View, CA
$207,000 - $290,000Hybrid

About The Position

About Tapestry Tapestry is Alphabet’s moonshot for the electric grid, operating at the convergence of energy infrastructure and advanced AI. Born at X (the innovation engine behind Waymo, Verily, and Google Brain), Tapestry builds computational and analytical platforms that make the world’s power grids visible, predictable, and resilient. We provide AI-driven planning and simulation tools that allow system operators, utilities, and planners globally to operate more efficiently and integrate clean energy at scale. Tapestry currently collaborates with key partners across the U.S., U.K., Chile, New Zealand, Australia, and Brazil. About the Role We are seeking a Staff Software Engineer / Technical Lead to co-anchor the technical direction, systems architecture, and engineering standards for our Infrastructure and Developer Productivity team. In this role, you will partner alongside another L6 Tech Lead and the Engineering Manager to set our multi-year platform roadmap. Our systems must process massive amounts of power-grid data, run heavy scientific simulations, and support emerging AI workflows. You will treat our internal infrastructure as a product—building reliable, self-service tools for our developers, automating our delivery pipelines, and creating secure environments where engineers and automated agents can build and test software quickly and safely.

Requirements

  • 8+ years of production experience in Infrastructure, Site Reliability Engineering, DevOps, or Developer Productivity/Platform Engineering.
  • 2+ years serving as a formal Tech Lead or Staff Engineer directing the technical roadmap, architectural designs, and execution for a multi-pod or multi-team engineering surface.
  • Advanced production-grade expertise with Kubernetes (GKE), container networking, and multi-tenant cloud architectures (Google Cloud Platform preferred).
  • Deep architectural experience with modern declarative tools (Terraform, Pulumi, or similar) managing complex multi-environment cloud footprints.
  • Proven track record building large-scale, automated build/test/release pipelines (e.g., Tekton, GitHub Actions, Argo Workflows, Bazel) designed around self-service internal developer platforms.
  • Demonstrated ability to interview internal engineering stakeholders, quantify developer friction points, and deliver platforms that measurably increase overall deployment frequency and reduce MTTR.

Nice To Haves

  • Practical experience designing infrastructure, sandboxes, and execution runtimes specifically geared toward AI agents, LLM evaluations, or high-performance GPU orchestration.
  • Hands-on architectural exposure to supporting large-scale data platforms (e.g., BigQuery, Spark, Kafka, Ray) or complex distributed simulation environments.
  • Working knowledge of Google-internal infrastructure primitives (Borg, Monarch, Spanner, Piper/Blaze) or experience operationalizing an X moonshot into an independent production environment.
  • Background engineering high-throughput platform tooling or developer infrastructure at companies operating at high engineering scale (e.g., Netflix, Snowflake, LinkedIn, Datadog).

Responsibilities

  • Architect and operate scalable, multi-tenant Kubernetes clusters on Google Cloud Platform (GCP). Ensure efficient compute scheduling and autoscaling across mixed hardware workloads (CPUs, GPUs, and TPUs) supporting simulation and machine learning models. Design resilient networking, IAM boundaries, and secure multi-project cloud environments.
  • Design fast, hermetic, and automated build, test, and release pipelines that reduce cycle times for product engineers. Provide reliable, on-demand testing environments so teams can validate changes safely before production. Build automated deployment and rollback mechanisms with clear canary verification.
  • Drive declarative, reproducible cloud infrastructure using modern Infrastructure-as-Code (such as Terraform). Implement automated policy checks, secret management, and container vulnerability scanning into standard deployment workflows.
  • Design secure, isolated sandbox environments that allow automated AI tools and agents to safely run tests and inspect code. Identify high-leverage opportunities to automate repetitive developer workflows using modern AI tools.
  • Partner with your fellow L6 Tech Lead to split architectural ownership, guide system designs, and run engineering design reviews. Mentor mid-level and senior engineers (L4/L5), raising the technical bar for code reviews, testing, and system design.
  • Define team-wide standards for metrics, logs, and distributed tracing to ensure high visibility into production health. Partner with data and ML teams to establish SLAs/SLOs, lead disaster recovery exercises, and run blameless post-mortems.

Benefits

  • Competitive salary and equity
  • Medical, dental, and vision coverage
  • Generous PTO and flexible hybrid work model
  • 401(k) with employer contribution
  • Professional development
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service