Forward Deployed AI Engineer

gravity9
Hybrid

About The Position

gravity9 is expanding its Forward Deployed Engineering team to build production agentic AI systems for enterprise clients, in partnership with leading frontier-model providers. We already support clients from different verticals to have agentic and RAG systems live in production across healthcare, financial services, retail and global logistics. As a Forward Deployed AI Engineer you work embedded in the client's environment, from discovery, through architecture and build, to production and handover. This is end-to-end agent engineering, not proofs of concept and not advisory work. Two flavours of engagement: Internal enterprise use cases, such as reconciliation and KYC workflows in financial services, clinical and claims workflows in healthcare, operations and supply-chain workflows in logistics. Product- and customer-facing agents, more greenfield, often exploratory, built into the client's own product. Engagements are typically a team of engineers over a few months, working shoulder-to-shoulder with the client's team including architects, DevOps and QA. Embed with the client. Work hand-in-hand inside the client's environment, codebase and cloud tenancy, often in a hybrid team alongside their engineers. You are visible to the client from day one. Design and build production agentic AI systems. Multi-agent orchestration, tool and function calling, retrieval, planning and routing, human-in-the-loop checkpoints, state and checkpointing, guardrails, failure handling and recovery. Do the unglamorous data work. A large share of every engagement is data engineering: ingestion, flattening deeply nested structures, extracting content from unstructured documents, classification, summarisation, tagging, enrichment, indexing. Models reason over data, bad data beats a good model every time. Own evaluation and accuracy. Establish a baseline eval dataset at the start of the engagement, automate grading, and track groundedness, faithfulness and retrieval quality per tool, not just at the agent level. Be ready to defend accuracy numbers to a sceptical enterprise stakeholder. Engineer for cost and latency. Model selection and routing (cheaper, faster models for non-reasoning steps; frontier models where reasoning genuinely earns it), prompt and context budgeting, caching. Cost is a non-negotiable metric on every engagement. Ship it properly. Observability and tracing, CI/CD, IaC, monitoring the client can actually operate, security and compliance review. Transfer knowledge deliberately. We don't run a long-term support business. Every engagement is designed so the client owns and can extend the system after we leave. You architect with their team in the room, pair with their engineers, and hand over working monitoring and documentation.

Requirements

  • Strong Python (async, typing, testing, packaging).
  • Hands-on production experience with frontier models: prompt and context engineering, structured outputs, tool and function calling, streaming, token and context-window management.
  • Built and shipped multi-agent or agentic workflows, planner/router patterns, supervisor and sub-agent designs, ReAct-style loops, state and checkpointing, retries and interrupts, human-in-the-loop review gates.
  • At least one of: Anthropic Agent SDK, LangGraph, LlamaIndex, or an equivalent orchestration framework, plus the judgement to know when to use none of them.
  • Chunking strategy, embeddings, hybrid and semantic search, re-ranking, citation and provenance, natural-language-to-query translation.
  • NoSQL databases such as MongoDB (aggregation pipelines, Atlas Search, Atlas Vector Search) or a strong equivalent, plus SQL.
  • Building ingestion and enrichment pipelines over messy structured and unstructured sources, documents, PDFs, object storage.
  • Building eval harnesses and eval datasets, LLM-as-judge with its limits understood, regression testing of prompts and agents, metric selection per use case.
  • Production delivery on AWS (incl. Bedrock), Azure or GCP, containers, serverless, networking basics, secrets management, IAM.
  • Git, code review, testing, CI/CD, IaC (Terraform or equivalent), observability.
  • Client presence and credibility.
  • Can hold a technical conversation with a client architect and a business conversation with their CTO or head of operations in the same hour, and be trusted by both.
  • Can present, whiteboard, and answer hard questions without deflecting.
  • Turns a business problem described by non-technical people such as a clinician, campaign manager or supply-chain planner into a technical design, and explains the technical design back in their language.
  • Contribute strongly, disagree well, and take direction without friction.
  • Make progress anyway, and make the ambiguity visible rather than hiding it.
  • Instinctively asks "how does this get deployed, monitored and maintained?" rather than stopping at a working notebook.
  • Actively enjoys upskilling the client's engineers, because self-sufficiency at handover is the definition of success, not follow-on billing.
  • Unblock yourself, chase the access request, and follow the thread to the answer.
  • Understands that scope, cost and the client's willingness to pay are part of the engineering problem, and contributes to scoping and estimating honestly.
  • New client, new domain, new stack every few months; occasional travel; occasionally a sceptical stakeholder who has been told AI is coming for their job. You stay steady and constructive.
  • Clear design docs, decision records, handover material and status updates in English.
  • Genuine curiosity about the field.
  • 3+ years software engineering, 1+ years hands-on LLM / agentic work with at least one system live in production.
  • Client-facing exposure.

Nice To Haves

  • Working competence in TypeScript / Node.
  • Anthropic certification (Claude Developer / Anthropic-issued credential).
  • Anthropic Agent SDK production experience.
  • MCP (Model Context Protocol): building servers and clients, tool exposure, auth patterns.
  • Claude Code as an autonomous SDLC agent: sub-agents, hooks, custom skills, agent marketplaces, context sharing across a team.
  • LLM observability and tracing tooling (LangFuse, LangSmith, Arize, Braintrust or similar).
  • Regulated-environment delivery: HIPAA, GDPR, SOC 2, FCA/PRA, PII handling, data residency, guardrails and red-teaming.
  • Graph or taxonomy-based knowledge representation alongside vector retrieval.
  • Kafka / streaming, Databricks, or comparable large-scale data platform experience.
  • Voice and multimodal agents; evaluation of non-text outputs.
  • FinOps for AI workloads; unit-economics modelling for agent systems.
  • Open-source contribution to the agent / LLM ecosystem, or conference speaking.
  • Prior experience as an FDE, solutions architect, or delivery consultant at a frontier-model, data platform, or infrastructure vendor.
  • 5+ years engineering, 2+ years AI or agentic AI, has owned the architecture of at least one production agent system end-to-end and led a client conversation about it.
  • 7+ years, has led delivery teams of 3–7 people, can act as engagement tech lead, contributes to pre-sales and estimation, and can mentor a growing FDE bench.

Responsibilities

  • Embed with the client.
  • Work hand-in-hand inside the client's environment, codebase and cloud tenancy, often in a hybrid team alongside their engineers.
  • Design and build production agentic AI systems.
  • Do the unglamorous data work.
  • Own evaluation and accuracy.
  • Engineer for cost and latency.
  • Ship it properly.
  • Transfer knowledge deliberately.

Benefits

  • Support for vendor certifications and access to partner training programmes.
  • Company-wide AI tooling as part of how we work day to day.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service