Senior AI Engineer, Agentic Systems

Career.io,
Remote

About The Position

We are developing a new product, an AI career companion app. It is distinct from our main platform and is in its early stages. At its heart is a conversational assistant which helps people with their job search, along with an agentic feature that enables it to carry out actions on their behalf such as updating a resume, searching for jobs, preparing for an interview, and following up on the outcome. You would have control of that core since the team is small and the direction has already been decided, most of the technical choices being yours. This is a 100% remote/work-from-home role.

Requirements

  • 5+ years as a software or ML engineer, including at least 2 years shipping LLM-based products to real users.
  • Experience building a conversational product that people used, with a clear account of successes and failures.
  • Production experience with agent frameworks (LangGraph or similar), tool calling, MCP, structured outputs, and state management.
  • Experience evaluating multi-turn systems: offline eval sets, judge validation, online testing.
  • Experience with guardrails and adversarial input in a live product.
  • Experience with streaming, latency, and cost optimization for LLM systems.
  • Proficiency in Python, asynchronous programming, and API design.
  • Experience with Git, Docker, and AWS.
  • Interest in the problem domain: careers, coaching, how people make decisions about work.

Nice To Haves

  • Fine-tuning or post-training on conversational data (SFT, DPO, distillation).
  • Voice interfaces: STT, TTS, real-time conversational agents.
  • Consumer product experience.
  • Background in coaching, education, health, or similar fields.
  • MCP server development.
  • Classical ML or recommendation systems.

Responsibilities

  • Develop a multi-turn assistant that maintains context across sessions over weeks or months.
  • Implement user memory and personalization, tracking what someone has shared, their progress, and changes.
  • Perform retrieval over the coaching corpus and fine-tune if beneficial.
  • Design conversation flows for when the assistant should ask, suggest, act, or refer the user elsewhere.
  • Orchestrate multi-step agent workflows with explicit state, tool calling, and handoffs between specialized agents.
  • Determine which actions require user confirmation and which do not.
  • Route requests between model providers based on cost, latency, and quality, with fallbacks.
  • Implement traces for every agent decision to facilitate debugging of failures.
  • Create offline evaluation sets for multi-turn conversations.
  • Utilize LLM-as-judge scoring, validated against human ratings.
  • Conduct online experiments on live traffic, measured against user outcomes (applications, interviews, offers).
  • Implement guardrails for scope, tone, and escalation in a consumer product.
  • Handle prompt injection and jailbreak attempts.
  • Ensure GDPR and EU AI Act compliance as part of the design.
  • Optimize for streaming responses and latency.
  • Manage cost per conversation through caching, routing, and prompt compression.
  • Deploy using FastAPI, async Python, Postgres, Redis, and AWS.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service