About The Position

Aircall is seeking a Machine Learning Engineer specializing in Evals and Voice Models to build out the evaluation foundation across their AI products. This role involves working on voice models, agent capability evaluations, benchmark design, LLM-as-judge systems, failure analysis, and the infrastructure to support these efforts. The goal is to establish shared metrics, test sets, and tooling to consistently measure accuracy, resolution quality, and safety across products. The engineer will set up repeatable pipelines for regression testing and benchmarking as models and features evolve, ensuring teams can ship confidently without reinventing evaluation methodology for each product.

Requirements

  • BS in Computer Science, Machine Learning, Statistics, or related field
  • 3+ years of experience in ML Engineering or Applied ML with 8+ years of overall experience
  • Strong experience in evaluating supervised, unsupervised, LLMs and deep learning models.
  • Hands-on experience in failure analysis and evaluating LLMs
  • Experience building automated evaluation systems
  • Strong communication skills to articulate complex technical concepts across technical and non-technical audiences
  • Hands-on experience training or fine-tuning voice/speech models (TTS, ASR, or speech-to-speech), including data pipeline construction and experimentation.

Nice To Haves

  • MS / PhD in Computer Science, Machine Learning, Statistics, or related field
  • Experience evaluating LLMs or agentic systems (e.g., LLM-as-a-judge, RAG evaluation)
  • Experience with synthetic data generation and prompt engineering
  • Experience training or fine-tuning voice models at scale, with familiarity in synthetic data generation, model distillation, or low-latency inference optimization for production voice agents.

Responsibilities

  • Design and document comprehensive evaluation frameworks for Aircall’s AI agents across voice, chat and messaging.
  • Train and fine-tune voice models (TTS, ASR, speech-to-speech) using production and synthetic data, iterating on architecture, data mix, and training strategy to improve accuracy, naturalness, and latency.
  • Assess AI generated solutions across training pipelines, experimentation setups, debugging processes, and optimization strategies.
  • Analyze system design decisions and identify strengths, weaknesses, and potential failure points.
  • Design annotation guidelines and workflows for human-labeled evaluation data, and calibrate LLM-as-judge systems against human raters to ensure automated evals stay trustworthy over time.
  • Build and maintain live quality monitoring for deployed AI agents, tracking accuracy, resolution rate, and safety signals in production, and flagging model or data drift before it impacts customers.
  • Own the metric contract for every published AI metrics, including definition, population, grain, rollup, validity window.
  • Build release gates, the offline regression suite each AI surface must pass before a prompt, model, or config change ships, measuring reliability across repeated trials, not just average pass rates.
  • Build voice-specific evaluation: simulated callers across accents, languages, background noise, barge-in, DTMF, and tool failures, with latency and ASR accuracy as first-class quality metrics.

Benefits

  • Competitive salary package & benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service