Senior ML Systems Engineer, Inference

Runpod
•$150,000 - $220,000•Remote

About The Position

Runpod is the AI Developer Cloud, serving over a million developers who use the platform to experiment, train, fine-tune, deploy, and scale AI. Having processed more than 20 billion inference requests and recently closing a $100M Series A, Runpod is at a pivotal moment in AI infrastructure. The company is seeking individuals who are passionate, driven, and aim to make a significant impact at scale. This is a remote-first role within a small, agile team that values ownership and speed. This role focuses on making Runpod the premier platform for LLM inference, prioritizing speed and cost-efficiency. The engineer will lead efforts to optimize LLM serving performance end-to-end, encompassing measurement, analysis, and improvement across various models, hardware, and workloads. The work will directly influence customer experience regarding latency and cost. It's a hands-on engineering position for someone adept at identifying and resolving performance bottlenecks, and implementing robust production solutions.

Requirements

  • 5+ years of professional system engineering experience.
  • Deep, hands-on experience with vLLM, SGLang (or a comparable serving engine) in production or at serious benchmark scale.
  • Strong software engineering skills in Python, with comfort in large, performance-critical codebases.
  • Solid understanding of LLM inference performance drivers: batching, memory, parallelism, and latency/throughput trade-offs.
  • Experience with modern inference optimization techniques (e.g., quantization, speculative decoding, distributed serving).
  • Rigor in benchmarking and performance analysis, including proficiency with GPU profiling tools.
  • Ability to clearly articulate results in writing and translate them into actionable decisions.

Nice To Haves

  • Experience writing or tuning GPU kernels in CUDA or Triton.
  • Contributions to inference or ML systems projects.
  • Experience with multi-node GPU systems and high-speed networking.
  • Experience at a company where inference cost and latency were core business metrics.

Responsibilities

  • Define inference performance metrics (throughput, time to first token, inter-token latency, cost per token) and develop rigorous, repeatable measurement tooling.
  • Profile and diagnose performance issues across the serving stack, from scheduling and memory management to kernels and interconnects.
  • Enhance serving efficiency for large, state-of-the-art models on single-node and multi-node GPU deployments.
  • Translate performance insights into production-ready runtimes, configurations, and defaults for automatic customer benefit.
  • Collaborate with product and infrastructure teams to shape inference offerings on Runpod.
  • Stay abreast of the rapidly evolving inference ecosystem, including the open-source community, to inform decisions on adoption, development, and contribution.
  • Identify and resolve bottlenecks in the serving engine/runtime when configuration tuning is insufficient.

Benefits

  • Competitive base pay ($150,000 - $220,000)
  • Meaningful equity (stock options)
  • Generous medical, dental & vision plans
  • Flexible PTO
  • Remote work first
  • $1,200 Home Office & Equipment Stipend
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service