About The Position

The enterprise software landscape is fracturing, and Zendesk is leading the shift from "Systems of Record" to "Systems of Action" with the Resolution Platform. This platform aims to be the first CX company to secure $1B in AI-driven revenue by solving the "Orchestration Trap." The role involves building a high-performance runtime that allows fleets of AI Agents to Perceive, Reason, and Act, operating at the cutting edge of large language models, planning algorithms, and multi-agent coordination. The successful candidate will design advanced memory structures and autonomous learning mechanisms to bridge the gap between offline model capabilities and dynamic, real-world task execution. Zendesk currently runs production AI agents that autonomously resolve customer service tickets across over 100,000 accounts, handling planning, execution through live APIs, and self-learning through synthesized resolution patterns. The role will focus on pushing this architecture further by addressing challenges in plan decomposition, memory management, selective skill acquisition, and multi-agent delegation. Additionally, it involves developing domain-specialized agent models trained via RL on production trajectories, building the RL training infrastructure, hardening evaluation processes with quality gates integrated into CI, and implementing enterprise-scale guardrails for autonomous agents.

Requirements

  • 5+ years building production ML/AI systems.
  • Hands-on experience in agent architectures (planning, tool dispatch, memory, error recovery).
  • Strong evaluation instincts and experience building internal evals to close the gap between public benchmarks and production performance.
  • Python and PyTorch fluency.
  • Familiarity with at least one agent framework, with the judgment to know when to build custom solutions.

Nice To Haves

  • Experience with or genuine depth in RL for language models: reward shaping, online/offline tradeoffs, reward hacking as a diagnostic signal.
  • Experience with LangChain tutorials is not sufficient for this role.

Responsibilities

  • Design and implement advanced memory structures and autonomous learning mechanisms for AI agents.
  • Develop and refine the iterative architecture for AI agents, focusing on plan decomposition, memory management, and skill acquisition.
  • Build and own the science and systems for training domain-specialized agent models using Reinforcement Learning (RL) on production trajectories.
  • Develop RL training infrastructure, including reward curricula and rollout systems.
  • Enhance the evaluation suite for AI agents, implementing multi-turn evaluation and automated trajectory analysis.
  • Integrate quality gates into CI to block deploys when agent performance drops.
  • Implement multi-layered defenses and guardrails for autonomous agents, including supervisor patterns, capabilities-based access control, and output validation.
  • Own both the science and systems for training domain-specialized models.

Benefits

  • Hybrid way of working
  • Fulfilling and inclusive experience
  • Flexibility to work remotely for part of the week
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service