About The Position

Mirantis is seeking a Senior DevOps Engineer specializing in Storage to deploy, integrate, and operate high-performance storage for GPU-accelerated compute and AI platforms. This role involves managing the storage layer where Kubernetes intersects with bare metal, including setting up NFS-based storage, integrating it into clusters via CSI, and optimizing it for demanding AI workloads. The work will encompass hybrid, edge, and air-gapped deployments built on the Mirantis K0rdent stack. The ideal candidate views storage as infrastructure to be automated, observed, and tuned, possesses strong Kubernetes storage expertise, deep Linux storage and networking fundamentals, and proficiency in infrastructure-as-code and GitOps practices. They should be self-directed in problem-solving, capable of setting operational standards, and effective in cross-team communication. While bare-metal hardware experience is a plus, deep Linux storage knowledge is essential.

Requirements

  • 7+ years of experience in SRE or infrastructure operations
  • 5+ years of building/operating distributed production Storage systems at scale
  • Hands-on with High Performance Storage solutions (VAST, Weka, DDN, PowerScale)
  • Linux and Kubernetes storage fundamentals (NFS, CSI)

Nice To Haves

  • Bare-metal hardware experience

Responsibilities

  • Integrate NFS-based high-performance storage (e.g., VAST, Dell PowerScale) into Kubernetes clusters via CSI, storage classes, and persistent volumes.
  • Tune the NFS data path — mount options, nconnect/RDMA, Linux client, and network settings — for high-throughput, low-latency GPU/AI workloads.
  • Deploy and operate storage services and operators; manage capacity, quotas, snapshots, and lifecycle.
  • Configure and optimize Linux systems for storage workloads, including driver setup, file system layout, network tuning, and kernel parameter optimization.
  • Deliver storage integration for k0s-based Kubernetes via Cluster API (CAPI) and K0rdent management/child cluster topologies.
  • Operate storage in fully disconnected (air-gapped) environments, including local artifact/mirror connectivity (Harbor) and PKI/TLS considerations.
  • Automate storage provisioning and configuration with infrastructure-as-code (Terraform/OpenTofu) and GitOps pipelines (ArgoCD or Flux).
  • Build monitoring, alerting, and observability for storage performance, capacity, and health.
  • Diagnose and resolve performance, reliability, and scaling issues across the storage stack.

Benefits

  • Professional development and training
  • Attend conferences and working groups
  • Company outings, happy hours, hackathons, and tech talks
  • Competitive compensation package with a strong benefits plan
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service