Gen AI Inferencing Engineer

Shree Narayani Networking SolutionsDallas, TX

About The Position

We are seeking a Gen AI Inferencing Engineer with a strong "Infra/MLOps-first" mindset. The ideal candidate will have a background in ML platform/infrastructure engineering, rather than primarily application development. Experience running models in production at scale is crucial, with a focus on deployment using vLLM or Triton Inference Server. This includes experience with containerization (Docker/K8s) and real-world throughput/latency tuning. Strong MLOps skills are essential, encompassing CI/CD for ML pipelines, fine-tuning workflows, and understanding inference framework internals. The role involves owning the infrastructure that other data science teams build upon, providing shared tooling rather than one-off notebooks. While knowledge of RAG is a plus, the primary focus is on efficient model serving rather than retrieval logic development. A good fit would be someone with an ML platform, SRE-for-ML, or MLOps background who has specific experience with GenAI serving, beyond traditional ML model serving.

Requirements

  • ML platform/infra engineering background
  • Experience running models in production at scale
  • Deployment experience with vLLM or Triton Inference Server
  • Experience with containerization (Docker/K8s)
  • Real throughput/latency tuning experience
  • Strong MLOps skills
  • CI/CD for ML pipelines
  • Fine-tuning workflows
  • Understanding of inference framework internals
  • Experience owning infrastructure for data science teams
  • Experience with GenAI serving

Nice To Haves

  • RAG knowledge

Responsibilities

  • Run models in production at scale
  • Deploy models via vLLM or Triton Inference Server
  • Containerize models using Docker/K8s
  • Perform real-world throughput and latency tuning
  • Implement CI/CD for ML pipelines
  • Manage fine-tuning workflows
  • Understand inference framework internals
  • Own and manage infrastructure for data science teams
  • Develop and maintain shared tooling
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service