The Infra Ops team is a new team in the Execute pillar of Weights & Biases infrastructure. Our mission is to keep our platform engineers building: we absorb, reroute, and automate incoming infrastructure requests — incidents, how-to questions, troubleshooting requests, and one-off asks — so that other pillars can focus on their deliverables. Build and release engineering is part of the Execute pillar — an organized, safe, and robust release process is the natural way to reduce post-release incidents, rollbacks, and toil. We work across the full W&B/CoreWeave infra stack — Kubernetes, Go, Terraform, ClickHouse, CircleCI, ArgoCD, Argo Rollouts, GitHub Actions, and others — running on GCP, AWS, and Azure, across multi-tenant SaaS, dedicated cloud, and on-prem deployments. We are seeking an Infrastructure Operations Engineer to be part of the first line of support for the infrastructure org. This is a support-oriented, interrupt-driven role: your days are shaped by incoming requests from our internal customers — Solutions Engineers, security, compliance, and product developers on other teams — rather than by a single long-running project. With guidance from senior teammates, you'll triage and resolve many requests end to end, and help convert recurring issues into documentation and automation to prevent repeats. It's a fast way to learn the entire infrastructure surface, and a natural "landing pad" into infrastructure engineering. You'll write code most days — automation and scripts in Python, Bash, and Go — and take part in a first-responder rotation covering the daily request peak in Slack and office hours. Your customers are primarily internal, though you'll occasionally work a customer escalation when our merchant-support and SA teams need help.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior