Senior Software Engineer, AI Agent & LLM

Otter.aiMountain View, CA
$185,000 - $230,000Hybrid

About The Position

We are looking for a Senior AI Agent & LLM Engineer who combines strong software engineering capabilities with a deep focus on AI quality. You will help build and improve the AI systems behind Otter’s conversational knowledge engine and AI Chat, spanning user-facing agent experiences, evaluation systems, and the shared platform required to operate them reliably at scale. This is a hands-on role for someone who can move from an ambiguous product problem to a working production solution, define how quality should be measured, and drive improvements across services, models, and user experience.

Requirements

  • 5+ years of AI Agent engineering, machine learning engineering, or related experience.
  • Strong backend or distributed-systems engineering skills.
  • Hands-on experience building and shipping AI-agent or LLM-powered products.
  • Strong focus on AI quality and experience evaluating nondeterministic systems.
  • Can use data, traces, logs, and qualitative examples to identify and resolve complex failures.
  • Works effectively across services, technical domains, and organizational boundaries.
  • Combines strong product judgment with rigorous engineering and evaluation practices.
  • Operates with high agency, strong ownership, and a bias toward action.
  • Can take ambiguous problems from initial exploration through production launch and continuous improvement.

Nice To Haves

  • Demonstrated ability to use coding agents effectively while rigorously reviewing and controlling the quality of their output is a plus.

Responsibilities

  • Build AI-agent experiences for Otter AI Chat that help users reason over conversations, retrieve knowledge, and complete complex tasks.
  • Advance Otter’s conversational knowledge engine through better knowledge extraction, annotation, and indexing.
  • Develop robust evaluation datasets, automated graders, regression tests, and release gates for AI quality.
  • Diagnose failures across models, prompts, retrieval, tools, data pipelines, backend services, and product workflows.
  • Build shared agent infrastructure for orchestration, tracing, debugging, retries, sandboxed execution, and observability.
  • Improve task completion, correctness, groundedness, reliability, latency, and cost.
  • Turn production traces, customer feedback, and usage signals into measurable product and model improvements.
  • Collaborate across AI, product, infrastructure, data, security, and application teams to deliver end-to-end capabilities.

Benefits

  • We provide reasonable accommodations for qualified applicants throughout the hiring process.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service