We're looking for a Senior AI/ML Test and Evaluation Engineer to build and operate the benchmarking and evaluation capability at the core of an AI platform. This is a role for someone who is more interested in what a model gets wrong than in what it gets right. You build the evaluation harnesses — automated metrics paired with structured human expert judgment, applied to candidate models and to the agentic workflows built on top of them. You develop repeatable methodologies for comparing performance against current operational baselines, which means the comparison holds up when someone runs it again in six months with a different model. And you document the limitations and surface the failure modes that matter, including the ones nobody asked about. Your reports go to senior stakeholders and inform decisions about which capabilities are ready to field. That's the weight of the job: a benchmark that looks good and hides a failure mode is worse than no benchmark at all, and you're the check against that. This is hands-on engineering on an open-source toolchain. This position is contingent upon contract award. Travel of up to 15% may be required, primarily to Government facilities and between company locations.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior