About The Position

The Custom Agent Service handles the full agent lifecycle: an admin configures an agent with instructions, knowledge articles, and actions through a UI; a customer ticket triggers execution; the agent plans, acts through real APIs via the Integration Action Platform, and resolves the issue. The backend is Python, the infrastructure is Kubernetes on AWS, and the agent architectures range from single-pass ReAct loops to our iterative multi-plan executor. We are onboarding our first internal and external EAP customers, so the work ships to real accounts with real tickets. We are looking for help with several key areas: Agent execution core: Improving the planning loop, tool dispatch, memory integration, and error recovery that make up the main execution path. This involves working directly on the code that determines the agent's next steps, aiming to enhance its speed, reliability, and capabilities. It also includes integrating new architectures into the production path and ensuring their robustness for real-world traffic. Knowledge retrieval: Optimizing the retrieval pipeline (embedding, reranking, context assembly) for customer knowledge bases at runtime. The goal is to balance answer quality with latency and token cost across thousands of heterogeneous knowledge bases per deployment. Actions and connectors: Enhancing the execution layer for agents calling Zendesk APIs, third-party connectors (Shopify, Salesforce, etc.), custom actions, and other agents via A2A. This includes ensuring reliable retries, timeouts, schema validation, and graceful degradation when connectors fail. It also involves self-servicing new connector integrations through the Connector SDK. Production instrumentation for model training: Instrumenting the execution pipeline to capture implicit reward signals (resolution success, escalation patterns, user feedback) that will feed into the ML team's training pipeline for domain-specialized models. Security and compliance: Developing PII filtering, audit logging, action versioning, and governance patterns to keep agents within admin-configured bounds, working directly with Product Security on security review items as the platform scales.

Requirements

  • 5+ years of backend engineering with strong Python skills.
  • Shipped production systems, not just models.
  • Understand the difference between getting an agent to work locally and running it across 100,000 accounts.
  • Comfortable across the full agent stack: LLM APIs, prompt engineering, tool calling, memory management, evaluation.
  • Can build an agent loop from scratch, and know when a framework helps vs. when it gets in the way.
  • Think about what happens when the model returns garbage, the connector times out, and the customer is waiting.
  • Build for the failure case, not just the happy path.

Responsibilities

  • Work directly on the code that decides what the agent does next and make it faster, more reliable, and more capable.
  • Integrate new architectures into the production path and harden them for real traffic.
  • Balance answer quality against latency and token cost in the retrieval pipeline.
  • Ensure reliable retries, timeouts, schema validation, and graceful degradation when connectors fail mid-execution.
  • Self-service new connector integrations through the Connector SDK.
  • Instrument the execution pipeline to capture implicit reward signals that feed into the ML team's training pipeline.
  • Develop PII filtering, audit logging, action versioning, and governance patterns.
  • Work directly with Product Security on security review items as the platform scales.
  • Ship working code, review PRs carefully, and communicate clearly about what is done, what is blocked, and what is at risk.

Benefits

  • Hybrid way of working, enables us to purposefully come together in person, at one of our many Zendesk offices around the world, to connect, collaborate and learn whilst also giving our people the flexibility to work remotely for part of the week.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service