Senior Staff Software Engineer, Cloud Availability Platform

Crusoe•Sunnyvale, CA
•$250,000 - $300,000

About The Position

Crusoe is building the operating system for the AI datacenter. They operate one of the world's largest managed GPU fleets and are looking for engineers to build a unified platform that senses, reasons about, and acts on the entire fleet, moving towards fleet autonomy. This is a platform team, meaning the focus is on building an API-first system with SDKs, paved paths, and a self-serve portal, treating internal teams as customers. The platform consists of four layers: agents for telemetry and execution, a distributed infra graph for system connections, a reconciliation core for state management, and domain services for various operational tasks. The team operates on a continuous autonomy loop (sense, correlate, reason, act, learn) and aims to automate recurring manual interventions.

Requirements

  • 10+ years building distributed systems, control planes, or infrastructure platforms.
  • Strong software engineering skills in Go, Python, or Rust.
  • Experience building platforms other engineers consume: public or internal APIs, SDKs, or developer tooling with real adoption.
  • Depth in at least one of: workflow/orchestration engines (Temporal or similar), event-driven architectures, graph data models, policy/rules engines, or reconciliation-based control loops (Kubernetes operator patterns).
  • Experience running what you build: you have carried a pager for a platform other teams depend on.
  • Systems thinking across the hardware/software boundary.

Nice To Haves

  • Internal developer platforms: API gateways, service catalogs, Backstage-style portals, micro frontend architectures.
  • GPU or bare-metal fleet infrastructure: DCGM, Redfish/IPMI, firmware lifecycle.
  • High-cardinality observability platforms (per-GPU telemetry at fleet scale).
  • InfiniBand or RoCE fabrics.
  • AI agents applied to infrastructure triage and autonomous remediation.

Responsibilities

  • Design and build core platform services including RBAC, tenancy, workflow engine, policy engine, and state reconciler.
  • Design the public face of the platform: API gateway, resource model, and versioned API contracts.
  • Build SDKs, workflow templates, golden paths, and the developer portal with micro frontend framework.
  • Build the inventory and topology graph and pipelines to ensure its accuracy against reality.
  • Build site, GPU, and network agents and the event bus for reliable telemetry and command movement.
  • Deliver the platform roadmap, including piloting site foundation, first site deployment, zero-downtime firmware upgrades, automated RMA, and scaling to 100K+ GPUs with specific performance targets (MTTD under 60 seconds, MTTR under 30 minutes).
  • Work with embedded engineers to translate operational needs into extensible services.

Benefits

  • Competitive compensation and equity packages
  • Restricted Stock Units
  • Paid time off, paid holidays & leave of absence programs
  • Comprehensive health, dental & vision insurance
  • Employer contributions to HSA account
  • Paid parental leave
  • Paid life insurance, short-term and long-term disability
  • Professional development & tuition reimbursement
  • Mental health & wellness support
  • Commuter benefits (parking & transit)
  • Cell phone stipend
  • 401(k) Retirement plan with company match up to 4% of salary
  • Volunteer time off
  • Global travel insurance & emergency assistance
  • Daily meals allowance
  • Additional perks & programs specific to location
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service