Applied AI/ML Engineer

OvertoneNew York City, NY
Onsite

About The Position

As a founding AI Engineer at Overtone, you will build and own the AI systems that power our core product experience, from conversational intelligence to matchmaking insight generation. You'll design prompting, evaluation, and model infrastructure that makes our AI reliable, trustworthy, and continuously improving. This is a hands-on, early-stage role: you'll run experiments, ship infrastructure, and build the tools that give the team real visibility into model behavior and quality.

Requirements

  • 5+ years of engineering experience, with meaningful time spent building LLM-powered products in production
  • Strong prompt engineering and evaluation instincts
  • Experience building evaluation pipelines for LLM systems
  • Familiarity with vector databases and retrieval-augmented architectures
  • Experience with fine-tuning techniques such as SFT, DPO, or RLHF
  • Strong understanding of when ML models outperform prompt-based systems
  • Experience working in early-stage environments with high ownership
  • Experience with Python, SQLAlchemy, FastAPI, PostgreSQL

Nice To Haves

  • Experience with recommender systems or ranking systems
  • Experience with MLOps and model lifecycle infrastructure
  • Experience building AI evaluation frameworks or model observability systems
  • Experience building internal developer tooling (React/TypeScript + Python)
  • Experience with production voice pipelines: streaming ASR (Deepgram, AssemblyAI, Whisper), TTS (ElevenLabs, Cartesia, PlayHT), or realtime voice agent frameworks
  • Experience with latency-sensitive or streaming AI surfaces, including WebRTC/LiveKit and evaluating voice quality (WER, naturalness, perceived responsiveness)

Responsibilities

  • Design and maintain the prompts, structured outputs, and orchestration systems that power Overtone's AI-driven features, including text-based LLM output and transcripts generated from our code voice AI experience.
  • Build evaluation infrastructure to measure AI quality at scale using tools such as Braintrust, Langfuse, TensorZero, and Promptfoo, ensuring LLM judges are aligned and evals cover generated text and voice transcripts, including transcription fidelity and response latency.
  • Run structured experiments across models, prompts, and configurations to optimize quality, cost, and latency.
  • Build internal tools (dashboards, admin interfaces, debugging workflows) that give the team visibility into system behavior and model performance.
  • Apply data science and lightweight modeling to improve match quality, identifying when traditional ML models or ranking systems outperform prompt-based approaches.
  • Own the model serving layer: deploying models, managing inference infrastructure (API providers and self-hosted), model versioning, and cost/latency optimization in production. Work with the backend lead to define clean service contracts between the AI layer and the rest of the system.
  • Identify opportunities to fine-tune open-source models using techniques such as SFT, DPO, and RLHF for cost and latency improvements.
  • Design and maintain vector databases and memory architectures that enable personalization and context-aware experiences.
  • Ensure systems are fair, unbiased, and auditable, building evaluation pipelines and human-in-the-loop processes where appropriate.
  • Work closely with product, engineering, and research to bring AI capabilities into real user-facing features with scalability and reliability in mind.

Benefits

  • Salary of $200k to $300k
  • Meaningful stock options with considerable upside potential
  • Health insurance
  • Dental insurance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service