About The Position

The enterprise software landscape is fracturing. We are transitioning from "Systems of Record" to "Systems of Action." Zendesk is leading this shift with the Resolution Platform, aiming to be the first CX company to secure $1B in AI-driven revenue. To achieve this, we must solve the "Orchestration Trap." We are not just building features; we are building the high-performance runtime that allows fleets of AI Agents to Perceive, Reason, and Act. You will operate at the cutting edge of large language models, planning algorithms, and multi-agent coordination. By designing advanced memory structures and autonomous learning mechanisms, you will bridge the gap between offline model capabilities and dynamic, real-world task execution.

Requirements

  • Deep foundation in machine learning, transformer architectures, and applied agentic systems.
  • History of designing real-world AI applications and leveraging agentic frameworks to build reliable, multi-step automated workflows.
  • Understand that building a great agent requires orchestrating specification, adaptive planning, tool execution, and iterative synthesis while embedding strict policies directly into the agent loop.
  • Know how to decompose high-level goals into actionable, verifiable sub-tasks.
  • Obsessed with the science of evaluation and know how to close the distribution mismatch between how an agent performs in a sandbox versus how it behaves when navigating the ambiguity of real-world production.
  • Python, PyTorch, and applied agent frameworks (e.g., LangChain, LangGraph, or similar orchestration tools).
  • Experience with custom LLM simulators, continuous multi-turn evaluation environments, and automated failure analysis pipelines.

Responsibilities

  • Building the World's Best Task Agents: Achieve state-of-the-art performance against the industry's most rigorous benchmarks. Optimize our agentic workflows to push past the 2026 frontiers, targeting elite-level autonomous problem solving on complex multi-step tool-use benchmarks.
  • Advanced Memory & Cognitive Architectures: Design sophisticated memory systems inspired by human cognition, allowing agents to dynamically filter interference, maintain context, and leverage long-term historical knowledge effectively across extended interactions.
  • Trajectory Analysis & Reasoning Refinement: Analyze complex agentic AI trajectories and trace patterns to understand how models navigate non-deterministic, multi-step problems. Build systems that capture the agent's entire chain-of-thought, analyzing conditional branches to actively detect anomalies and halt hallucination loops prior to failure.
  • Self-Improving Loops & Skill Discovery: Architect scalable, autonomous self-improving loops that allow agents to operate and learn continuously without human intervention. Design frameworks where tool search patterns, error handling, and sophisticated retry logics are actively fed back into the system to dynamically improve the structure of the agentic AI planner. Enable agents to autonomously discover, synthesize, and incrementally acquire new reusable skills based on environmental feedback and task completion.
  • Enterprise Guardrails & Content Safety: Engineer multi-layered defenses to secure agentic workflows against unique risks such as tool misuse, cascading action chains, and unintended control amplification. Design strict input validation to block malicious prompt injections or jailbreak attempts, as well as output filtering to ensure responses remain within the application's domain boundary.
  • Supervisor Patterns & Governance: Implement governance-centric architectures, such as supervisor or manager agent patterns, to explicitly regulate actions and decision sequences during runtime. Enforce capabilities-based access, ensuring that an agent's available tools are strictly determined by the user's role, and align system evaluations with compliance standards like the NIST AI Risk Management Framework.
  • Rigorous Agentic Evaluation: Move beyond static leaderboards to build continuous, multi-turn evaluation frameworks. Design preference data pipelines to rigorously test emergent multi-agent coordination, ensuring our agents act safely and align perfectly with enterprise intent.

Benefits

  • Fulfilling and inclusive experience
  • Hybrid way of working
  • Flexibility to work remotely for part of the week
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service