Aircall is seeking a Machine Learning Engineer specializing in Evals and Voice Models to build out the evaluation foundation across their AI products. This role involves working on voice models, agent capability evaluations, benchmark design, LLM-as-judge systems, failure analysis, and the infrastructure to support these efforts. The goal is to establish shared metrics, test sets, and tooling to consistently measure accuracy, resolution quality, and safety across products. The engineer will set up repeatable pipelines for regression testing and benchmarking as models and features evolve, ensuring teams can ship confidently without reinventing evaluation methodology for each product.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior