Staff Software Engineer (Cloud Infrastructure)

CrusoeSan Francisco, CA
$215,000 - $260,000Hybrid

About The Position

Crusoe is seeking a highly skilled and motivated GPU Fleet Operations Engineer to join Crusoe’s Fleet Operations team. This role is focused on the advanced diagnosis, maintenance, and repair of high-performance GPU compute clusters, ensuring maximum uptime, reliability, and performance across our fleet. The ideal candidate will be hands-on with GPU rack-level troubleshooting and work closely with data center operations, engineering, and vendors to support cutting-edge infrastructure featuring the latest NVIDIA and AMD GPUs. This position plays a critical role in maintaining the health and scalability of Crusoe’s rapidly growing GPU fleet.

Requirements

  • Ability to code in Golang
  • Proven experience diagnosing and repairing high-density, rack-mounted compute hardware in production environments.
  • Deep understanding of GPU architectures and hands-on experience with GPU-based systems.
  • Experience supporting NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X series platforms.
  • Familiarity with high-speed interconnects such as InfiniBand, NVLink, and RDMA over Converged Ethernet (RoCE).
  • Strong Linux experience (Ubuntu, Rocky Linux, CentOS) using the command line for diagnostics and testing.
  • Proficiency with GPU and system diagnostic tools such as NVIDIA DCGM and NVIDIA field diagnostic utilities.
  • Experience working with enterprise server hardware, power delivery, and cooling systems.
  • Strong analytical and problem-solving skills.
  • Excellent communication and collaboration skills.
  • Ability to work independently in a fast-paced data center or operations environment.

Nice To Haves

  • Technical certification or Associate’s/Bachelor’s degree in Electrical Engineering, Computer Science, or a related field or demonstrated experience.
  • Experience working directly with hardware vendors and escalations.
  • Background in large-scale GPU fleet operations or hyperscale data center environments.

Responsibilities

  • Automate deep-level diagnosis and troubleshooting of hardware faults within GPU racks and high-density compute systems.
  • Develop software to troubleshoot and support GPU platforms including NVIDIA A100, H200, GB200, B200, B300 and AMD 350X / 355X.
  • Execute component-level diagnosis and remediation for failed or degraded hardware.
  • Partner with data center operations to manage and perform field-replaceable unit (FRU) repairs for GPUs, power supplies, cooling systems, interconnects, and networking hardware.
  • Conduct post-repair validation, burn-in testing, torch testing, and NVIDIA NCCL testing to ensure system stability and performance.
  • Implement and execute preventative maintenance procedures to improve fleet reliability and extend hardware lifespan.
  • Perform firmware and BIOS upgrades across the GPU fleet.
  • Maintain detailed documentation of maintenance activities, failures, and resolutions in ticketing and asset management systems.
  • Develop and update standard operating procedures (SOPs) for troubleshooting, repair, and validation workflows.
  • Collaborate with engineering, software, and data center operations teams to identify root causes of systemic failures and implement preventative solutions.
  • Participate in a rotating infrastructure on-call schedule (about one week every 4–6 weeks) with daytime coverage and handoff to the Europe team.

Benefits

  • Hybrid work schedule
  • Industry competitive pay
  • Restricted Stock Units in a fast growing, well-funded technology company
  • Health insurance package options that include HDHP and PPO, vision, and dental for you and your dependents
  • Employer contributions to HSA accounts
  • Paid Parental Leave
  • Paid life insurance, short-term and long-term disability
  • Teladoc
  • 401(k) with a 100% match up to 4% of salary
  • Generous paid time off and holiday schedule
  • Cell phone reimbursement
  • Tuition reimbursement
  • Subscription to the Calm app
  • MetLife Legal
  • Company paid commuter benefit; $300 per pay period
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service