Senior Software Engineer, Infrastructure Operations - Weights & Biases

Weights & Biases•Sunnyvale, CA
•$153,000 - $204,000•Hybrid

About The Position

The Infra Ops team is a new team in the Execute pillar of Weights & Biases infrastructure. Our mission is to keep our platform engineers building: we absorb, reroute, and automate incoming infrastructure requests — incidents, how-to questions, troubleshooting requests, and one-off asks — so that other pillars can focus on their deliverables. Build and release engineering is part of the Execute pillar — an organized, safe, and robust release process is the natural way to reduce post-release incidents, rollbacks, and toil. We work across the full W&B/CoreWeave infra stack — Kubernetes, Go, Terraform, ClickHouse, CircleCI, ArgoCD, Argo Rollouts, GitHub Actions, and others — running on GCP, AWS, and Azure, across multi-tenant SaaS, dedicated cloud, and on-prem deployments. We are seeking an Infrastructure Operations Engineer to be part of the first line of support for the infrastructure org. This is a support-oriented, interrupt-driven role: your days are shaped by incoming requests from our internal customers — Solutions Engineers, security, compliance, and product developers on other teams — rather than by a single long-running project. With guidance from senior teammates, you'll triage and resolve many requests end to end, and help convert recurring issues into documentation and automation to prevent repeats. It's a fast way to learn the entire infrastructure surface, and a natural "landing pad" into infrastructure engineering. You'll write code most days — automation and scripts in Python, Bash, and Go — and take part in a first-responder rotation covering the daily request peak in Slack and office hours. Your customers are primarily internal, though you'll occasionally work a customer escalation when our merchant-support and SA teams need help.

Requirements

  • Bachelor’s degree or foreign equivalent in Computer Science, Computer Engineering, Data Science, Information Systems, or a closely related technical or quantitative field involving substantial coursework in Engineering.
  • 2+ years in an infrastructure, SRE, operations, DevOps, or technical-support role — or equivalent experience from a computer-science (or related) degree, a bootcamp, or substantial personal/open-source projects. We hire on demonstrated ability, not credentials.
  • Comfortable scripting in at least one of Python, Bash, or Go, and eager to use code to eliminate repetitive work.
  • Experience with Linux, containers, and at least one major public cloud (GCP, AWS, or Azure); infrastructure-as-code (Terraform) is a must.
  • Able to drive well-scoped problems to resolution, and to recognize when to pull in a more senior teammate.
  • Strong written and verbal communication and a service mindset — you communicate effectively with our internal customers.
  • A genuine desire to help people and unblock them — you take satisfaction in solving someone else's problem, not just your own.
  • Patience and composure under pressure, with the ability to stay positive and constructive.
  • Comfortable with a support-driven, high-context-switching, and sometimes repetitive workload — you can hold several small threads at once without losing the plot, and you drive the automation of recurring patterns.
  • Aim to make yourself unnecessary — automate everything to free time for larger infrastructure projects at the edges of other infrastructure pillars. We have much ground to cover.

Nice To Haves

  • Exposure to CI/CD systems (GitHub Actions, Argo, or similar) and observability tooling (Datadog).
  • Applied AI: Hands-on experience designing, building, and maintaining complex, production-grade AI agents — including continuous multi-source data ingestion and indexing, and RAG / context augmentation.
  • Hands-on experience with production databases (PostgreSQL, ClickHouse, or similar).
  • Prior internship or role on a platform, developer-experience, or support-engineering team.

Responsibilities

  • Triage and resolve incoming infrastructure requests across Slack, Jira, and office hours — resolving well-scoped ones yourself and escalating others with a clear, reproducible hand-off.
  • Troubleshoot problems across the stack — dig through logs, systems, and configuration to work out what is actually happening, asking for help when a problem runs deep.
  • Write scripts and other automation in Python, Bash, and Go that reduce repetitive work.
  • Help maintain and improve build and release pipelines, along with the runbooks and self-service docs the team depends on.
  • Partner and collaborate closely with our SA, security, and product engineers on one side, and with the other infra pillars on the other.
  • Take part in a 24/7 escalation on-call rotation
  • Document recurring how-to and configuration questions so they can be answered once and reused.

Benefits

  • Medical, dental, and vision insurance - 100% paid for by CoreWeave
  • Company-paid Life Insurance
  • Voluntary supplemental life insurance
  • Short and long-term disability insurance
  • Flexible Spending Account
  • Health Savings Account
  • Tuition Reimbursement
  • Ability to Participate in Employee Stock Purchase Program (ESPP)
  • Mental Wellness Benefits through Spring Health
  • Family-Forming support provided by Carrot
  • Paid Parental Leave
  • Flexible, full-service childcare support with Kinside
  • 401(k) with a generous employer match
  • Flexible PTO
  • Catered lunch each day in our office and data center locations
  • A casual work environment
  • A work culture focused on innovative disruption
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service