Senior AI Engineer

DV Trading
$200,000 - $300,000Onsite

About The Position

DV Trading is building a centralized AI function and is now hiring for the model layer. The long-term goal is for DV to own its model capability — not to be permanently dependent on what frontier providers choose to offer, at what price, for how long. This role is how that happens: fine-tuning and distilling open-weight models for DV-specific tasks, operating the inference infrastructure to run them on-prem, and building the model gateway that routes intelligently across open and closed providers. The near-term result is lower cost and better latency. The long-term result is a firm that controls its own AI stack.

Requirements

  • 5+ years software engineering; strong Python
  • Production fine-tuning or distillation of open-weight models (not just inference API wrappers)
  • Experience serving LLMs on-prem (vLLM, TGI, Triton, or equivalent)
  • Experience managing GPU infrastructure (provisioning, scheduling, utilization monitoring) in a production environment
  • Model evaluation and regression testing in production
  • Kubernetes and GPU workload management
  • Strong grasp of the tradeoffs between open and closed models across cost, quality, latency, and data sensitivity

Nice To Haves

  • Quantization, PEFT/LoRA, or other efficient training techniques
  • Model gateway or inference proxy design (routing, fallback, rate limiting)
  • Financial services or other regulated/sensitive-data environments
  • Familiarity with the open model ecosystem (Hugging Face, model cards, licensing)

Responsibilities

  • Build and operate a model gateway routing inference across open and closed models with cost, latency, and quality tracking
  • Design and run distillation pipelines: use frontier model outputs to generate training data for task-specific open models
  • Fine-tune and evaluate open-weight models (Llama, Qwen, Mistral, or similar) for DV-specific tasks
  • Deploy and maintain on-prem inference infrastructure (vLLM, TGI, or equivalent) on Kubernetes
  • Build model evaluation frameworks for quality, cost, latency, and regression
  • Define criteria and tooling for model selection: when open models are production-ready vs. when to use closed APIs
  • Partner with the agent engineering team to ensure the model layer meets agent workload

Benefits

  • Discretionary bonus eligibility
  • Medical, dental, and vision insurance
  • HSA, FSA, and Dependent Care Options
  • Employer Paid Group Term Life and AD&D insurance
  • Voluntary LTD, Life & AD&D insurance
  • Flexible Vacation policy
  • Retirement plan with employer match
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service