Senior Production Engineer, Storage

CrusoeSunnyvale, CA
$170,000 - $205,000

About The Position

At Crusoe Energy Systems, our Site Reliability Engineering (SRE) team plays a mission-critical role in maintaining the performance and reliability of our AI-optimized cloud infrastructure. The Storage-focused SRE role is responsible for ensuring the availability, performance, and scalability of Crusoe's cloud storage products and services, which power compute-intensive, latency-sensitive workloads for AI and HPC use cases. This role directly supports our vertically integrated, sustainable cloud platform by building and optimizing distributed, fault-tolerant storage systems at scale.

Requirements

  • Bachelor's degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience
  • 5+ years of professional experience in Storage SRE, systems, or storage engineering
  • Deep, hands-on experience with enterprise storage platforms such as Pure Storage or EMC — not limited to volume creation/provisioning, but a working understanding of how the systems are architected and operated
  • Deep understanding of object, block, and file storage paradigms
  • Proficiency in a programming language such as Go, Python, Java, or C
  • Experience with Infrastructure as Code and deployment tooling such as Terraform, Ansible, or Puppet
  • Deep knowledge of Linux internals with a focus on I/O subsystems, memory management, and storage scheduling
  • Familiarity with storage protocols like NFS, SMB, iSCSI, or NVMe-oF
  • Strong experience working with containerized workloads and orchestration platforms (e.g., Kubernetes, Docker)
  • Excellent incident response, troubleshooting, and documentation practices
  • Experience with building and operating managed services at scale such as object, file and block storage (AWS, GCP, Azure)

Nice To Haves

  • Hands-on experience with distributed storage systems (e.g., Ceph, GlusterFS, OpenEBS, Vast, Lightbits)
  • Contributions to open-source storage projects or the Linux storage stack
  • Experience with hybrid storage models across on-prem and cloud environments

Responsibilities

  • Build automation and self-healing tools to monitor and maintain Crusoe's distributed cloud storage infrastructure, which includes block, file, and object storage systems.
  • Drive reliability initiatives focused on data replication, encryption, backup and restore strategies, and robust failover mechanisms.
  • Collaborate closely with storage engineers to implement and maintain high-performance NVMe- and SSD-backed volumes that support large-scale AI compute clusters.
  • Support user-facing storage services with a focus on availability, performance tuning, and adherence to error budgets.
  • Investigate and resolve storage-related incidents using deep telemetry, logs, and performance profiling.
  • Partner with hardware and kernel teams to diagnose low-level I/O issues and optimize I/O paths, cache policies, and file systems.
  • Contribute to the architecture of fault-tolerant, scalable storage backends tailored for AI-first cloud environments.

Benefits

  • Industry competitive pay
  • Restricted Stock Units in a fast growing, well-funded technology company
  • Health insurance package options that include HDHP and PPO, vision, and dental for you and your dependents
  • Employer contributions to HSA accounts
  • Paid Parental Leave
  • Paid life insurance, short-term and long-term disability
  • Teladoc
  • 401(k) with a 100% match up to 4% of salary
  • Generous paid time off and holiday schedule
  • Cell phone reimbursement
  • Tuition reimbursement
  • Subscription to the Calm app
  • MetLife Legal
  • Company paid commuter benefit; $300 per month
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service