AI Infrastructure & Platform Operations Engineer

Mirantis

1d•Remote

About The Position

Mirantis is seeking an AI Infrastructure & Platform Operations Engineer to join their European AI Infrastructure & Platform Operations team. This team is responsible for operating large-scale AI infrastructure environments powered by NVIDIA GPUs, high-performance networking, Kubernetes, and next-generation platform technologies. The role involves ensuring the availability, performance, and operational stability of critical AI infrastructure platforms deployed across multiple datacenters. The engineer will work at the intersection of infrastructure, networking, and platform operations, supporting environments that power modern AI workloads and contributing to the evolution of AI-powered operational services through platforms like k0rdent AI.

Requirements

3+ years of experience in infrastructure operations, platform operations, network operations, site reliability engineering, cloud operations, datacenter operations, or related technical roles.
Strong Linux administration and troubleshooting skills.
Good understanding of networking concepts and experience diagnosing infrastructure-related issues.
Working knowledge of Kubernetes in production environments.
Experience supporting production infrastructure and services.
Strong analytical and problem-solving skills.
Experience working within structured operational and incident management processes.
Excellent communication and collaboration skills.
Ability to work within a shift-based operational environment.

Nice To Haves

Experience in NVIDIA GPU infrastructure and accelerated computing platforms.
InfiniBand networking and NVIDIA UFM.
Kubernetes platform operations.
AI infrastructure or HPC environments.
Site Reliability Engineering (SRE) or Platform Engineering.
Observability platforms such as Grafana, Prometheus, ELK, or OpenTelemetry.
Infrastructure automation technologies and Infrastructure-as-Code practices.
Large-scale distributed systems and production platforms.

Responsibilities

Monitor, operate, and support production AI infrastructure platforms.
Investigate and resolve infrastructure, networking, hardware, and platform-related incidents.
Support NVIDIA GPU infrastructure and associated platform services.
Monitor and troubleshoot Kubernetes-based environments.
Investigate performance, availability, and reliability issues across infrastructure and platform components.
Collaborate with engineering teams, hardware vendors, datacenter personnel, and service delivery teams to resolve technical issues.
Participate in incident response, root cause analysis, and operational improvement activities.
Contribute to improvements in monitoring, observability, automation, and operational processes.
Maintain operational documentation, runbooks, and knowledge articles.

Benefits

Work with some of the most advanced AI infrastructure environments in production today.
Gain exposure to NVIDIA GPU technologies, Kubernetes platforms, and high-performance networking environments.
Help define how next-generation AI infrastructure is operated and supported.
Be part of a team shaping the future of AI-powered operations through k0rdent AI.
Join a growing organisation investing heavily in AI infrastructure and platform services.

Stand Out From the Crowd

Upload your resume and get instant feedback on how well it matches this job.

Upload and Match Resume