Senior Inference Engineer

vCluster LabsNew York, NY
$190,000 - $230,000Remote

About The Position

As a Senior Inference Engineer at vCluster Labs, you are the first engineer we're hiring to own inference. You'll partner directly with our CTO to build the platform's inference layer from the ground up, taking models and turning them into a production-grade, query-to-response pipeline running at scale. From there, you will help lead the engineering direction of inference at vCluster, partnering with Product to shape what we build next as the space evolves.

Requirements

  • Production LLM serving experience: You've deployed and served LLMs using vLLM, SGLang, or TensorRT-LLM, ideally at a company built around inference at scale.
  • Inference optimization know-how: Hands-on experience with quantization, batching, caching, and routing, not just familiarity with the terms.
  • Hands-on programming experience: Strong engineering skills in Python or Golang, with real production code experience.
  • Communication: Strong communication skills, explaining technical concepts clearly to both engineers and non-technical stakeholders.

Nice To Haves

  • Familiarity with containerized environments (Docker, Kubernetes)
  • Hands-on generative AI experience with common ML frameworks (PyTorch, Transformers)
  • Good understanding of the GPU stack: CUDA, NCCL, drivers, and related libraries
  • Knowledge of model architectures and fine-tuning approaches
  • Experience with NVIDIA Dynamo

Responsibilities

  • Deploying models to production: Take LLMs and put them into production across one or more machines on GPU infrastructure, owning the full pipeline from a customer's query to the served response.
  • Serving frameworks: Stand up and operate serving infrastructure using vLLM, SGLang, or TensorRT-LLM.
  • Optimizing for scale: Apply quantization, batching, caching, and routing to keep latency and cost in check as traffic grows.
  • Programming, not just configuring: Build real infrastructure in Python or Golang — this is an engineering role, not a research or data-science one.
  • Owning the roadmap: Build the first iteration alongside our CTO, then take the lead on the inference platform and partner with Product to decide what we build next.

Benefits

  • Competitive Salary: We offer a competitive compensation package, including equity.
  • Platinum-Level Insurance: Health, dental, vision, and life Insurance, including plans for you and eligible dependents (benefits vary depending on country).
  • Flexible Working Schedule: You have a doctor’s appointment or need to head to the supermarket to get groceries at 2pm? We won’t have an issue with that. To us, results matter more than clocking in and out at the same time every day.
  • Workplace Flexibility: We’re very flexible about where you work. We know things can change in life and we’re happy to adjust the work environment for you along the way.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service