Member of Technical Staff - ML Infra

Trajectory•San Francisco, CA

About The Position

As an ML Infrastructure Engineer at Trajectory, you will build the infrastructure for AI systems that learn to improve their own training, inference, and kernels. Our goal is to unlock a 10× improvement every month somewhere in the stack - from GPU scale and model size to throughput, memory efficiency, caching, and latency. This role spans training, inference, and kernels. Bring deep expertise in at least one area and curiosity across the stack. We’ll shape your initial ownership around your strengths.

Requirements

  • Strong fundamentals in distributed systems, networking, storage, and failure recovery, with experience shipping and operating demanding systems.
  • Deep specialization in training infrastructure, inference systems, or GPU kernels, supported by systems built or measurable optimizations delivered.
  • Strong Python skills and languages relevant to your specialty, such as C++, CUDA, or Triton.
  • Understanding of PyTorch, JAX, or comparable framework internals, with strong profiling and debugging skills.
  • Ownership and clear communication: work closely with researchers and deliver measurable performance gains while preserving correctness and reliability.
  • Demonstrated capability over credentials.

Nice To Haves

  • Hands-on experience with training and RL stacks such as Miles, SkyRL, Prime Intellect’s verifiers, or comparable systems.
  • Depending on your specialty, experience with vLLM, SGLang, collective communication, or ML compilers is also valuable.

Responsibilities

  • Build and optimize distributed training and RL infrastructure, including rollout execution, GPU scheduling, checkpointing, and recovery.
  • Improve training experimentation throughput and shorten research iteration cycles.
  • Optimize serving for production agents and training rollouts.
  • Improve batching, scheduling, and KV-cache management while balancing latency, throughput, cost, and model quality.
  • Develop and optimize GPU kernels and runtimes using CUDA, Triton, or comparable tools.
  • Improve memory use and execution efficiency, preserving numerical correctness and verifying gains in real workloads.
  • Build reproducible benchmarks, observability, and automated research workflows that propose changes, run experiments, and validate improvements.
  • Work with researchers to turn new algorithms into reliable systems across training, inference, and kernels.
  • Build one of the world’s best continual learning loops for ML infrastructure: agents propose optimizations, run experiments, measure gains, and learn from the results across training, inference, and kernels.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service