Staff Applied AI Engineer, Product & Agent Performance

Arcadia
$175,000 - $200,000Remote

About The Position

Arcadia is seeking a Staff Applied AI Engineer, Product & Agent Performance to own product-layer decisions shaping agent behavior, including prompting, retrieval, context, memory, state, evaluation, and escalation. This role is crucial for ensuring Arcadia's agentic capabilities perform accurately, transparently, and safely at scale for clinicians, care teams, and patients. The successful candidate will partner with Product and Engineering to build systems that support responsible AI, enabling Arcadia to make evidence-based launch decisions and scale AI that is steerable, trustworthy, and ready for healthcare workflows. This is a staff-level individual contributor role focused on improving the performance, reliability, and cost-effectiveness of AI agents in a production environment.

Requirements

  • Equivalent practical experience demonstrating the depth required for this staff-level role.
  • 8+ years of production software engineering experience.
  • 3+ years of hands-on ownership of ML, LLM, or agentic systems in production.
  • Direct experience in healthcare, finance, or another regulated industry.
  • Demonstrated ability to diagnose agent failures, correctly attribute fixes to instruction, retrieval, context, or memory design, and weigh failures by severity and cost.
  • Hands-on experience with RAG architecture, production-grounded evaluation frameworks, and fallback or human-in-the-loop logic for automated systems.
  • Working familiarity with AWS AI/ML services, including Bedrock and SageMaker, sufficient to build and evaluate effectively in Arcadia’s environment.
  • Evidence-led judgment and the credibility to push back on launch decisions.
  • A builder’s instinct to run experiments and move from production failure to a fix.

Nice To Haves

  • Experience applying AI to healthcare data or workflows where safety, transparency, and calibrated uncertainty directly affect care teams or patients.
  • Experience with long-horizon, multi-turn or multi-agent workflows.
  • Experience with product-level AI documentation such as model cards.

Responsibilities

  • Design and iterate on agent behavior across real, live workflows, including long-horizon, multi-turn agentic tasks.
  • Design retrieval and context architecture to ensure the right source data reaches a model in the right structure, keeping agents grounded in real data.
  • Design memory and state handling across multi-turn and multi-agent flows, determining what information is carried forward, summarized, or dropped.
  • Create context and prompt templates that combine few-shot examples, structured formatting, and reasoning scaffolding for consistent agent behavior.
  • Improve performance through prompting, tool-use strategy, and context construction, validated through direct experimentation.
  • Build and run evaluations against real production conditions to measure performance, regressions, failure modes, and edge cases.
  • Author evaluation rubrics, quality heuristics, and thresholds that weight failures by severity and cost, and monitor these measures against production behavior.
  • Design and validate escalation paths that route agents to human review based on confidence and uncertainty, preserving safety and consistency under adversarial and edge-case conditions.
  • Design for cost-aware performance alongside latency, reliability, and accuracy through efficient context construction and tool-call economy.
  • Evaluate and sign off on model changes by baselining current behavior, running comparative evaluations, and making go/no-go decisions before a change reaches a customer.
  • Maintain product-level AI documentation, including model cards, intended use, limitations, and known failure modes.
  • Partner closely with Product and product managers to ensure agents are capable, steerable, trustworthy, and ready to scale.

Benefits

  • The opportunity to define how agent performance, safety, and readiness are measured for production healthcare workflows.
  • Meaningful ownership across prompts, context, memory, evaluations, and escalation patterns at product scale.
  • A cross-functional role translating production evidence into AI improvements used across Arcadia’s platform.
  • A mission-driven company working to improve how patients receive care.
  • A flexible, remote-friendly culture with personality and heart.
  • Employee-driven programs and initiatives for personal and professional development.
  • Membership in the talented, energized, diverse, and purpose-driven Arcadian community.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service