Staff Software Engineer, MetalDev

CoreWeaveSunnyvale, CA
$207,000 - $275,000

About The Position

As a Staff Software Engineer within our Compute Architecture organization, you will help build the software systems that operate the backbone of our large-scale GPU data centers. The METALDEV team builds Go-based distributed services that bring new infrastructure online, manage hardware lifecycle workflows, monitor production health, and automate safe operations across fleets of GPU servers and rack-scale systems. This is a software-first role at the intersection of distributed systems, production reliability, and hardware-aware automation, where your work directly improves the reliability, safety, and scalability of real-world infrastructure.

Requirements

  • B.S., M.S., or PhD in Computer Science or related field, or equivalent experience.
  • 8+ years of software engineering experience with a strong focus on infrastructure, cloud engineering, and distributed databases—particularly within large-scale datacenter and cloud environments.
  • Expertise in Go and proven experience building REST/gRPC APIs for mission-critical platforms.
  • Strong background in architecting and scaling cloud-native Kubernetes infrastructure and distributed services.
  • Proven success in mentoring engineers, leading technical projects, and influencing engineering strategy across teams.
  • Experience contributing to and collaborating with open source communities.
  • Skilled in applying a data-driven approach to reliability, optimization, and continuous improvement.
  • Excellent communicator able to work effectively with both technical and non-technical stakeholders.
  • Hands-on experience with observability stacks (Prometheus, Grafana, PromQL), CI/CD pipelines, and operating large fleets of GPU servers.
  • Track record of leading incident response, postmortems, and driving robust service reliability.

Nice To Haves

  • Working knowledge of Kafka, ClickHouse and CRDB.
  • DMTF, RedFish APIs, and GPU servers.

Responsibilities

  • Design, build, and operate Go-based services that manage the lifecycle of large-scale GPU data center infrastructure.
  • Build automation for data center bring-up, hardware discovery, health monitoring, remediation, and production operations.
  • Develop reliable APIs, services, and workflows for managing BMCs, firmware state, server health, and rack-level infrastructure.
  • Improve observability, alerting, and operational tooling so production issues can be detected, understood, and resolved quickly.
  • Translate incidents and hardware failure modes into software improvements that make the platform more resilient.
  • Partner with hardware-adjacent, infrastructure, operations, and software teams to design systems that work safely at fleet scale.
  • Provide technical leadership through design reviews, code reviews, architectural guidance, and mentorship.
  • Make pragmatic architecture decisions that balance reliability, simplicity, scalability, and operational burden.

Benefits

  • Medical, dental, and vision insurance - 100% paid for by CoreWeave
  • Company-paid Life Insurance
  • Voluntary supplemental life insurance
  • Short and long-term disability insurance
  • Flexible Spending Account
  • Health Savings Account
  • Tuition Reimbursement
  • Ability to Participate in Employee Stock Purchase Program (ESPP)
  • Mental Wellness Benefits through Spring Health
  • Family-Forming support provided by Carrot
  • Paid Parental Leave
  • Flexible, full-service childcare support with Kinside
  • 401(k) with a generous employer match
  • Flexible PTO
  • Catered lunch each day in our office and data center locations
  • A casual work environment
  • A work culture focused on innovative disruption
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service