Cloud Platform Software Engineer

Groq•New York, NY
•Hybrid

About The Position

Groq is building a global cloud platform powering production AI workloads, pioneered by the LPU—the first processor designed specifically for AI inference. The mission for this role is to own and evolve the Kubernetes platform and adjacent hyperscaler technologies that underpin Groq's cloud infrastructure. The engineer will build reliable, scalable platform capabilities that enable engineering teams to deploy and operate workloads efficiently as Groq's infrastructure grows. This role is based in the Dallas, San Francisco, or New York City area and will initially offer remote work flexibility, transitioning to onsite once local offices are established.

Requirements

  • Strong software engineering fundamentals with experience building and operating production infrastructure.
  • Deep hands-on experience with Kubernetes and proficiency designing Kubernetes-based architectures.
  • Experience owning Kubernetes environments in production, including deployment, scaling, upgrades, reliability, and troubleshooting.
  • Experience building software and automation for cloud or infrastructure platforms.
  • Strong understanding of distributed systems and the reliability and scalability challenges of operating infrastructure at scale.
  • Experience with cloud infrastructure and modern infrastructure-as-code and automation practices.
  • Ability to debug complex problems across application, Kubernetes, networking, compute, and infrastructure layers.
  • Comfortable taking end-to-end ownership of production systems and making pragmatic architectural tradeoffs.
  • Strong collaborator who can partner across engineering teams and translate their needs into scalable platform solutions.
  • Motivated by ownership, technical excellence, reliability, and building durable infrastructure.

Responsibilities

  • Own the architecture, development, reliability, and evolution of Groq's Kubernetes-based platform.
  • Design and build Kubernetes-based systems that operate reliably and efficiently at scale.
  • Build software, automation, and platform capabilities that simplify how workloads are deployed, operated, and scaled.
  • Improve the resilience, observability, performance, and operational efficiency of the Kubernetes platform.
  • Identify and eliminate reliability and scalability bottlenecks across clusters and supporting infrastructure.
  • Own platform capabilities through their full lifecycle, from design and implementation through production operations and continuous improvement.
  • Partner with software and infrastructure engineers to understand workload requirements and translate them into durable platform capabilities.
  • Establish and improve Kubernetes engineering patterns, tooling, and operational practices.
  • Debug complex production issues spanning Kubernetes, distributed systems, networking, compute, and cloud infrastructure.
  • Help shape the technical direction of the platform as its scale and requirements evolve.

Benefits

  • Long-Term Incentive (LTI) Program
  • Competitive compensation through Total Cash philosophy (inclusive of potential bonus value)
  • Robust suite of employee benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service