About The Position

Mirantis is seeking an experienced Software Engineer, Infrastructure to design and implement the Infrastructure Services that power their GPU-as-a-Service platform. The role involves building the control plane responsible for transforming high-level API calls into actual infrastructure actions, including enrolling bare-metal servers, provisioning them, and assembling them into multi-tenant Kubernetes clusters on high-performance hardware. The engineer will manage the entire lifecycle of infrastructure-level services, from Server and MachineType APIs to provisioning workflows and reconciliation loops that ensure the platform's hardware view aligns with physical reality.

Requirements

  • Strong experience designing RESTful APIs or gRPC services, including API versioning and gateway patterns.
  • Proficiency in Go.
  • Deep understanding of Kubernetes primitives and controller/reconciler patterns.
  • Hands-on experience with bare-metal provisioning flows (BMC/Redfish, PXE/iPXE, image management, hardware inspection).
  • Experience building workflow-driven or event-driven systems (e.g., Temporal) for long-running operations requiring state consistency across retries and failures.

Nice To Haves

  • Experience building informers or reconciliation bridges that keep an external datastore consistent with Kubernetes resource state.
  • Experience building platforms with strict data and network isolation between tenants.
  • Familiarity with Terraform/OpenTofu and GitOps-driven configuration (ArgoCD or Flux).
  • Familiarity with GPU server hardware, DPUs/NICs, and high-performance datacenter fabrics.

Responsibilities

  • Design, build, and maintain versioned REST and gRPC APIs for bare-metal server lifecycle, MachineType definitions, and cluster CRUD operations.
  • Develop asynchronous workflows for server enrollment, inspection, OS provisioning, and cluster bring-up, providing durable status updates.
  • Implement and maintain consoles and interfaces for visualizing hardware inventory, provisioning progress, and cluster health.
  • Design error-handling models, idempotency guarantees, and reconciliation loops for reliable management of long-running provisioning operations.

Benefits

  • Professional development and training
  • Attend conferences and working groups
  • Company outings, happy hours, hackathons, and tech talks
  • Competitive compensation package with a strong benefits plan
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service