Staff Platform Engineer

INFINITE CHOICE•Dallas, TX
•Remote

About The Position

InfiniteChoice is building its next-generation platform on Google Cloud. As our Staff Platform Engineer, you'll be the senior technical owner of that infrastructure: the GKE clusters, the Terraform that defines them, and the path traffic takes from the edge to our services. You'll also build what comes next: a self-service platform that lets our developers get infrastructure without filing a ticket, and the practices that keep our cloud footprint efficient as we grow. You'll report to the Director of Engineering, CloudOps, on a small team. Staff here means you set the technical direction for our infrastructure and raise the level of the engineers around you, while still writing a good share of the code yourself.

Requirements

  • 10+ years of experience in platform, SRE or infrastructure engineering, including ownership of production systems serving millions of requests a day.
  • Deep hands-on experience running Kubernetes in production, ideally GKE.
  • Strong Terraform experience, including structuring modules and state across many environments and GCP projects.
  • Solid GCP experience across networking, IAM, and core services.
  • You've led a migration of production workloads between platforms without downtime.
  • You write production code in Go, Python, or a similar language, and you can read application code when debugging.
  • You explain technical tradeoffs clearly in writing, to engineers and to non-engineers.
  • Bachelor's degree in Computer Science, Engineering, or equivalent professional experience

Nice To Haves

  • Industry certifications (Google Cloud Professional, SRE or related certifications preferred)
  • Practical experience with AI coding tools, and a view on where they help in infrastructure work and where they don't
  • Cloudflare, including WAF and bot management
  • Redis at scale, including clustered deployments
  • OpenTelemetry and Prometheus
  • Building or running an on-call and incident program

Responsibilities

  • Design and run the platform our product teams deploy onto, including templates, CI/CD, and GKE Gateway API routing, so teams can ship without filing a ticket with us.
  • Define SLOs with the teams that own each service, and use them to decide where reliability work goes.
  • Make architecture decisions for our infrastructure and write them up as ADRs others can review.
  • Review infrastructure changes across engineering and mentor the engineers making them.
  • Plan capacity for clusters, databases, and caches, find and remove waste, and keep GCP spend in line with traffic as we grow.
  • Take ownership of our GKE estate and Terraform, and get us to where every production infrastructure change goes through code review and Atlantis.
  • Bring structure to a GCP organization that grew quickly: project layout, group-based IAM, and cost visibility.
  • Ship the first self-service workflows for our developers, starting with the request types that fill most of our ticket queue today.
  • Turn our disaster recovery design into something we test on a schedule.
  • Participate in the on-call rotation and drive continuous improvements to incident response, operational readiness, and post-incident follow-through.

Benefits

  • Full autonomy to define processes, select technologies, and establish best practices
  • Direct impact on platform reliability serving millions of users
  • Opportunity to create lasting engineering culture and operational excellence
  • Remote-first culture with in-person meeting in Dallas, TX on need basis
  • Collaborative environment with smart, passionate engineers and cross-functional teams
  • Access to cutting-edge technologies and AI-driven development tools
  • Competitive compensation, equity participation, and comprehensive benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service