The Custom Agent Service handles the full agent lifecycle: an admin configures an agent with instructions, knowledge articles, and actions through a UI; a customer ticket triggers execution; the agent plans, acts through real APIs via the Integration Action Platform, and resolves the issue. The backend is Python, the infrastructure is Kubernetes on AWS, and the agent architectures range from single-pass ReAct loops to our iterative multi-plan executor. We are onboarding our first internal and external EAP customers, so the work ships to real accounts with real tickets. We are looking for help with several key areas: Agent execution core: Improving the planning loop, tool dispatch, memory integration, and error recovery that make up the main execution path. This involves working directly on the code that determines the agent's next steps, aiming to enhance its speed, reliability, and capabilities. It also includes integrating new architectures into the production path and ensuring their robustness for real-world traffic. Knowledge retrieval: Optimizing the retrieval pipeline (embedding, reranking, context assembly) for customer knowledge bases at runtime. The goal is to balance answer quality with latency and token cost across thousands of heterogeneous knowledge bases per deployment. Actions and connectors: Enhancing the execution layer for agents calling Zendesk APIs, third-party connectors (Shopify, Salesforce, etc.), custom actions, and other agents via A2A. This includes ensuring reliable retries, timeouts, schema validation, and graceful degradation when connectors fail. It also involves self-servicing new connector integrations through the Connector SDK. Production instrumentation for model training: Instrumenting the execution pipeline to capture implicit reward signals (resolution success, escalation patterns, user feedback) that will feed into the ML team's training pipeline for domain-specialized models. Security and compliance: Developing PII filtering, audit logging, action versioning, and governance patterns to keep agents within admin-configured bounds, working directly with Product Security on security review items as the platform scales.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed