AI Infrastructure Engineer

Mirantis
Remote

About The Position

We are looking for an AI Infrastructure Engineer to join our global team responsible for managing and supporting large-scale AI infrastructure environments. You will help ensure the availability, performance, and operational stability of critical AI infrastructure platforms, working closely with a distributed team across regions to provide continuous coverage and support. This is an opportunity to work hands-on with some of the most advanced Kubernetes-based AI infrastructure in production today, while contributing to the platforms and processes that keep it running reliably at scale.

Requirements

  • Proven experience managing and operating large-scale production systems (bare-metal and/or cloud).
  • Solid working knowledge of Kubernetes with excellent, demonstrable troubleshooting skills.
  • Experience configuring, customizing, and extending logging and monitoring tools (e.g., Prometheus, Grafana, ELK, or similar).
  • Experience with infrastructure automation technologies and Infrastructure-as-Code practices (e.g., Ansible, Terraform, or similar).
  • Effective verbal and written communication skills in English.
  • Strong analytical and problem-solving skills, with the ability to work through complex, ambiguous technical issues.
  • Willingness to occasionally work weekends and holidays.

Nice To Haves

  • Previous experience building, scaling, and running High-Performance Computing (HPC) environments.
  • Hands-on experiences with managing large scale Kubernetes platforms in production.
  • A good understanding of NVIDIA GPU technologies and the associated software stack.
  • Proficiency in scripting languages (e.g., Python, Bash, Go).

Responsibilities

  • Manage and operate production AI infrastructure environments.
  • Lead incident response and troubleshooting efforts and deliver timely service restoration during outages or performance degradations .
  • Troubleshoot infrastructure and networking issues across bare-metal and/or cloud environments with multiple vendors.
  • Conduct root cause analysis and drive product and operational improvements.
  • Contribute and improve operational documentation and knowledge base.
  • Collaborate with global team members across time zones to ensure continuous operational coverage, including occasional work during weekends and holidays.

Benefits

  • Professional development and training
  • Attend conferences and working groups
  • Company outings, happy hours, hackathons, and tech talks
  • Receive a competitive compensation package with a strong benefits plan
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service