Senior Agentic AI Engineer — Healthcare AI

Deloitte•Jersey City, NJ
•$110,700 - $372,900•Hybrid

About The Position

Deloitte has a new AI-first effort, backed by $1B in committed investment, building the reasoning models and agentic systems to rebuild how the healthcare system decides — across payers, providers, and life sciences, and for the patients they serve — so that care is faster, fairer, and far less wasteful. This is not AI applied at the margins. It is a ground-up rebuild of the decision-making machinery behind American healthcare, at national scale. This is an early, well-funded build. You will own agent systems end to end — from architecture through production — and your work ships into live clinical and operational settings within your first months, not into a lab. As a Senior Agentic AI Engineer, you will design, build, and operationalize the LLM- and SLM-powered systems behind real healthcare decisioning — the reasoning, orchestration, retrieval, memory, and control layers that let intelligent agents operate reliably across the hardest decisions in the industry: clinical reasoning, prior authorization and claims integrity, care navigation, and the operational workflows that run across payers, providers, and life sciences. This is not a prompt-only role. We are looking for builders who think deeply about system behavior, grounding, and reliability where a wrong action has real consequences for patients and the clinicians who serve them. You do not need a healthcare background. We pair every engineer with clinical and domain experts and teach you the domain — you bring the agentic engineering depth. We hire on demonstrated depth, not years — the level you join at is determined through our interview process, based on the depth and judgment you demonstrate, not your years in a title.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, Data Science, Computational Linguistics, or a related field.
  • Demonstrated depth building and shipping production agentic systems — this is your primary craft, not a recent exploration. We weigh shipped systems, research, model releases, and open source over years in a title; expect strong software/ML fundamentals plus substantial, recent hands-on agentic work.
  • Strong, hands-on experience building production agent systems with modern orchestration — LangGraph/LangChain or equivalent, including custom orchestration.
  • Experience designing and optimizing end-to-end RAG systems: indexing, retrieval, reranking, grounding, and evaluation.
  • Strong understanding of memory and context management, including context windows, retrieval-driven context assembly, persistent memory, and high-signal context selection.
  • Deep, practical understanding of LLM behavior — strengths, limitations, hallucination risks, reasoning constraints, and latency/cost trade-offs — and the evaluation methods used to measure them.
  • Experience evaluating and debugging agent behavior — task-success and trajectory analysis, not just output quality.
  • Strong Python engineering skills and modern software practices: testing, CI/CD, version control, and API integration; experience implementing observability, tracing, and debugging for LLM-based systems in production.
  • Hands-on experience with at least one frontier model platform (e.g., Anthropic, Google, OpenAI) and/or open-weight/self-hosted models (e.g., Llama via vLLM), including production tool use and agent capabilities.
  • Ability to travel 0–50%, on average, based on the work you do and the clients and industries/sectors you serve.
  • Limited immigration sponsorship may be available.

Nice To Haves

  • Experience with multi-agent systems and agent collaboration patterns.
  • Familiarity with vector databases and retrieval infrastructure such as Pinecone, Weaviate, or Milvus.
  • Exposure to model adaptation and fine-tuning techniques such as LoRA or QLoRA.
  • Understanding of traditional NLP concepts: tokenization, semantic similarity, entity extraction, summarization, and transformer fundamentals.
  • Experience operating in highly regulated, high-stakes, or operationally complex environments; healthcare exposure — clinical, payer, or life-sciences workflows, or standards such as FHIR — is a plus, not a requirement.
  • Demonstrated habit of staying current with AI research, benchmarks, and emerging engineering patterns.

Responsibilities

  • Design and implement agentic systems capable of multi-step reasoning, planning, tool use, and workflow execution against complex, regulated operational processes.
  • Build stateful workflows using frameworks such as LangGraph and LangChain — including branching, retries, self-correction, human-in-the-loop checkpoints, and reusable orchestration patterns.
  • Engineer for long-horizon reliability — multi-step task completion, recovery from compounding errors, planning under uncertainty, and robust tool use when individual steps fail.
  • Build the reasoning behind regulated decisions — policy- and criteria-grounded outputs, structured proposer/critic/judge-style review, and auditable rationales for high-stakes decisions across the industry, from clinical review and prior authorization to claims integrity and care management.
  • Develop end-to-end Retrieval-Augmented Generation (RAG) pipelines: ingestion, chunking, embeddings, vector and hybrid retrieval, reranking, contextual compression, and grounding strategies.
  • Engineer memory and context management — conversational state, persistent memory, retrieval-aware context assembly, and token-efficient context selection.
  • Apply modern context-delivery patterns (e.g., MCP-style tool/context interfaces) so agents access the right information at the right time.
  • Implement observability and tracing for prompts, tool calls, retrieval quality, agent traces, failures, drift, latency, and production behavior.
  • Apply guardrails, safety controls, and failure-handling to reduce hallucinations and unsafe actions.
  • Evaluate agents at the trajectory and task level — multi-step task success, failure-mode and regression analysis, and sandboxed test environments — alongside retrieval- and generation-quality metrics, automated checks, and human review.
  • Engineer healthcare-grade safety — deployment eval gates, human-oversight and escalation models, auditability and traceability for regulated decisions, and PHI/HIPAA-aware data handling.
  • Build integrations with internal and external tools, APIs, enterprise systems, databases, and model providers so agents operate safely within real business workflows.
  • Deliver production-quality code with strong practices in testing, CI/CD, logging, versioning, and documentation; make architecture decisions that balance quality, safety, latency, cost, and model risk.
  • Partner with our modeling and post-training engineers to improve model behavior for tool use, grounding, and long-horizon reasoning — through evaluation-driven feedback and, where it helps, fine-tuned or reasoning-optimized models.
  • Translate ambiguous, high-complexity operational processes into robust system logic and reusable AI patterns; stay current with advances in agentic systems and translate research into practical engineering decisions.

Benefits

  • Professional development opportunities
  • Mentorship programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service