About The Position

As a Senior / Staff Machine Learning Engineer in Wayve's AV Core organisation, you will advance reinforcement learning methods for end-to-end driving models. You'll identify where learning from reward or feedback can improve beyond behavior cloning, then take promising ideas from design through large-scale experiments, rigorous evaluation, and integration into our best driving models.

Requirements

  • You have a strong track record developing and experimentally validating reinforcement learning or closely related sequential decision-making methods on complex, high-dimensional problems.
  • You have a deep understanding of modern reinforcement learning fundamentals, including policy and value learning, off-policy learning, function approximation, distribution shift, and the failure modes of learned objectives.
  • You have hands-on experience with behaviour cloning, reinforcement learning, or related methods.
  • You're proficient in Python and PyTorch, with strong software engineering practices and hands-on experience building reliable machine learning training and evaluation systems.
  • You have excellent experimental judgement: able to turn an ambiguous behavioral problem into falsifiable hypotheses, useful metrics, disciplined ablations, and clear technical decisions.
  • You bring senior-level ownership and collaboration: able to lead a substantial technical area, work across research and engineering boundaries, and bring others along through clear written and verbal communication.

Nice To Haves

  • Experience with offline reinforcement learning, imitation learning, reward modeling, preference learning, or post-training of large neural policies.
  • Experience in autonomous vehicles, robotics, control, or another domain where policies interact with safety-critical physical systems, including an understanding of motion planning, vehicle dynamics, control, or collision avoidance.
  • Experience with closed-loop simulation, off-policy evaluation, uncertainty or calibration, and evaluation under rare or shifted conditions.
  • Experience training multimodal, transformer-based, or generative policy models at scale.
  • Proficiency in C++, CUDA, distributed training, or performance optimization for production machine learning systems.

Responsibilities

  • Shape and execute the reinforcement learning roadmap for Driving Core / Core Model Safety, selecting problems and methods against clear behavioral gaps and measurable success criteria.
  • Develop and evaluate post-behavior-cloning optimization methods, including offline and off-policy reinforcement learning as well as other reward-guided approaches; design the regularization, data strategy, and diagnostics needed to make policies reliably better.
  • Help improve the reward models and related learning signals used to train and evaluate driving policies, working with partner teams to strengthen their quality, scalability, and downstream usefulness.
  • Build robust training and experimentation workflows using large-scale driving data; diagnose distribution shift, objective misspecification, optimization instability, and data or evaluation bias.
  • Define evidence across offline metrics, open-loop tests, closed-loop simulation, and on-road evaluation, and distinguish genuine policy improvement from benchmark overfitting.
  • Productionize successful methods in the shared ML stack, communicate decisions and results clearly, and raise the technical bar through design reviews, code reviews, and mentoring.

Benefits

  • Salaries benchmarked against the market annually
  • Meaningful equity, sharing in the ownership and long term success of Wayve
  • Relocation support and visa sponsorship where applicable
  • Hybrid working, core hours and the chance to work hands on in vehicle workshops and labs
  • Learning and development budgets with support for training, conferences and growth
  • Comprehensive benefits including health insurance, dental, enhanced maternity and paternity leave, retirement or pension where applicable, access to therapists, wellbeing partnerships, team socials and more
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service