We are seeking a Gen AI Inferencing Engineer with a strong "Infra/MLOps-first" mindset. The ideal candidate will have a background in ML platform/infrastructure engineering, rather than primarily application development. Experience running models in production at scale is crucial, with a focus on deployment using vLLM or Triton Inference Server. This includes experience with containerization (Docker/K8s) and real-world throughput/latency tuning. Strong MLOps skills are essential, encompassing CI/CD for ML pipelines, fine-tuning workflows, and understanding inference framework internals. The role involves owning the infrastructure that other data science teams build upon, providing shared tooling rather than one-off notebooks. While knowledge of RAG is a plus, the primary focus is on efficient model serving rather than retrieval logic development. A good fit would be someone with an ML platform, SRE-for-ML, or MLOps background who has specific experience with GenAI serving, beyond traditional ML model serving.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed