HPC & MLOps Engineer

C-Gen.AIMountain View, CA
Remote

About The Position

At C-Gen.AI, we are pioneering the next generation of AI infrastructure. As a leading technology company, our mission is to deliver world-class AI Infrastructure management solutions that drive innovation, efficiency, and performance at scale. We seek passionate, highly skilled engineers who excel in dynamic environments and are eager to work with cutting-edge technologies to transform the computational landscape. Role Summary As an HPC & MLOps Engineer, you will be a foundational member of our Data Center team—integral to the design, deployment, and maintenance of high-performance computing (HPC) and MLOps systems. In this role, you will ensure our clients have access to reliable, scalable, and secure HPC environments. You will collaborate closely with cross-functional teams and work under a dynamic, startup environment that rewards initiative and technical excellence.

Requirements

  • 3+ years hands-on in HPC operations; strong familiarity with Slurm (or similar batch systems such as PBS Pro/OpenPBS).
  • Solid Python (including asyncio /task-oriented patterns) and Bash for automation, tooling, and data handling.
  • Deep understanding of batch queues, job lifecycle, and multi-tenant cluster policies.
  • Strong understanding of RedHat/Debian linux flavors and system administration.
  • Proficiency with at least one major cloud (AWS/GCP/Azure/OCI) and interest in others; comfortable bridging cloud and on-premises deployments.
  • Practical knowledge of Lustre/Ceph/NFS/object storage and the throughput/latency trade-offs common to HPC.
  • Understanding of AI training/inference workflows, GPU scheduling, drivers, and runtime management.
  • Confidence with high-performance networking (e.g., RDMA/InfiniBand/RoCE) and compiling Linux modules for network support.
  • Clear written/spoken English; able to collaborate across time zones and functions.
  • Self-directed, detail-oriented, and comfortable in a fast-moving startup.

Nice To Haves

  • Python libraries & SDKs: asyncio, aiohttp, cloud SDKs (e.g., boto3, google-cloud, azure-sdk, OCI).
  • CUDA/NCCL profiling, MPI optimization, kernel/sysctl tuning, GPU/CPU/IO benchmarking.
  • Experience balancing cost, performance, and reliability—especially in cloud bursting scenarios.
  • Having C/C++ and systems programming experience is a big plus.

Responsibilities

  • Configure, deploy, monitor, and maintain C-Gen.AI Cluster solution for a diverse client base.
  • Manage both cloud and on-premise deployments, ensuring optimal job scheduling and resource allocation.
  • Troubleshoot and optimize HPC library stacks—including OpenMPI, CUDA, TensorFlow, and PyTorch—and manage parallel file systems (e.g. Lustre, BeeGFS, Ceph, NFS, or object storage).
  • Develop and oversee automated deployments across multiple platforms (Cloud providers or on-premise clusters).
  • Implement best practices for network configuration, security, and cost optimization tailored to HPC needs.
  • Create and maintain Bash/Python scripts that streamline workflows, gather essential metrics, and empower self-service HPC management.
  • Contribute to the development and maintenance of HPC workflows for AI/ML teams, enhancing training and inference pipelines.
  • Collaborate with our in-house HPC experts on advanced projects involving performance tuning, MPI, and GPU parallelization.
  • Instrument systems, collect metrics, and build dashboards/alerts that enable self-service and rapid incident response.
  • Work closely with product, support, and customer teams; document designs, runbooks, and standards

Benefits

  • Competitive salary
  • stock options
  • flexible remote work arrangements
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service