Founding Senior Applied AI Engineer (Agentic Systems)

Puntt.aiSan Francisco, CA
Onsite

About The Position

Our mission is to free humanity from meaningless work. We are building the System of Record for Enterprise Compliance, turning hours of manual, high-stakes marketing workflows into minutes of automated precision. We are a small, high-density team of engineers and operators. Having proven our value with global brands, we are now at the inflection point where our technical architecture meets massive scale. This is a "rocket ship" moment: we are moving beyond simple automation into a world of vision-first, agentic workflows that solve the problems generic frontier models cannot. As Senior Applied AI Engineer, you’ll own the technical moat that makes us stand alone: the safety dataset and deterministic orchestration that prevents enterprise brands from ever trusting generic AI with compliance. You’re not building ‘better automation’—you’re building the category-defining infrastructure that gives us a 3-year lead in a market where second place doesn’t exist. You will lead the technical evolution of our agent-driven architecture: defining the role of each agent, designing multi-step reasoning flows, integrating tools, memory, and retrieval systems, and optimizing for accuracy, determinism, and trust. Your work will directly determine whether enterprise customers can rely on Puntt for legal and brand compliance at scale. This is a hands-on, in-office role in San Francisco, working closely with a small, senior team to build systems where correctness matters more than demos.

Requirements

  • 5–7+ years of professional engineering experience, with a strong record of shipping production systems
  • 1–2+ years building with LLMs in real applications (not just experimentation)
  • Expert Python experience
  • Hands-on experience designing RAG systems, vector search, embeddings, and structured retrieval
  • Strong engineering foundation with understanding of core CS concepts—data structures, concurrency, failure modes, and tradeoffs.
  • Experience building real systems with LLMs, not just experimenting with them.
  • Comfortable designing systems that combine LLMs with RAG, memory, graph-based context, and external tools rather than relying on a single prompt.
  • See multi-agent orchestration as a distributed systems challenge—latency, retries, observability, and consistency all matter.
  • Thrive in an early-stage environment where problems are underspecified and the best solution doesn’t exist yet.

Nice To Haves

  • Experience with LLM orchestration frameworks (e.g., LangGraph, CrewAI, or custom orchestration layers)
  • Experience with stateful workflow orchestration (Temporal a plus)
  • Experience operating AI systems on AWS (Lambda, S3, Bedrock, SageMaker, etc.)
  • Strong TypeScript experience
  • Bonus: experience with OCR, document parsing, or VLMs

Responsibilities

  • Design and evolve multi-agent LLM systems that decompose complex review tasks into reliable, auditable steps.
  • Define agent responsibilities, hand-offs, and termination conditions to minimize reasoning drift and maximize consistency.
  • Architect retrieval pipelines using RAG, structured memory, and emerging approaches like graph-based retrieval to provide agents with the right context at the right time.
  • Balance recall, precision, and latency across large knowledge bases (brand guidelines, regulations, historical decisions).
  • Own long-running, fault-tolerant workflows using Temporal (or similar), ensuring retries, versioning, and determinism across non-deterministic model calls.
  • Treat agent orchestration as a distributed systems problem: managing state, failures, and observability.
  • Build evaluation frameworks that go beyond “it looks right,” using statistical metrics, gold labels, and automated regression testing to prove system reliability.
  • Prioritize correctness and trust, especially in high-risk legal and compliance scenarios.
  • Collaborate on image and document preprocessing (OCR, layout analysis, VLMs) to ensure downstream agents receive structured, machine-readable context.
  • Focus on practical understanding, not computer vision research.
  • Move fluidly between Python-based LLM services, retrieval pipelines, and AWS infrastructure to ship reliable systems end-to-end.

Benefits

  • Foundational technical leader shaping how the system works
  • High-impact, high-trust domain
  • Speed without chaos
  • Ship quickly
  • Care deeply about system design, evaluation, and long-term reliability
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service