Senior AI/ML Test and Evaluation Engineer

OpenTeams•United States - Remote OR Hybrid, OR
•$145,000 - $250,000•Hybrid

About The Position

We're looking for a Senior AI/ML Test and Evaluation Engineer to build and operate the benchmarking and evaluation capability at the core of an AI platform. This is a role for someone who is more interested in what a model gets wrong than in what it gets right. You build the evaluation harnesses — automated metrics paired with structured human expert judgment, applied to candidate models and to the agentic workflows built on top of them. You develop repeatable methodologies for comparing performance against current operational baselines, which means the comparison holds up when someone runs it again in six months with a different model. And you document the limitations and surface the failure modes that matter, including the ones nobody asked about. Your reports go to senior stakeholders and inform decisions about which capabilities are ready to field. That's the weight of the job: a benchmark that looks good and hides a failure mode is worse than no benchmark at all, and you're the check against that. This is hands-on engineering on an open-source toolchain.

Requirements

  • Proven experience building and deploying software platforms with complex integration surface, preferably handling Machine Learning or agentic workloads.
  • Deep experience using python, including API frameworks and agentic harnesses.
  • High level knowledge of Computer vision, language modeling, and/or task specific agent evaluation.
  • Hands on experience with common frameworks and tooling, eg PyTorch and Hugging Face ecosystem.
  • Strong communication skills and experience collaborating effectively with cross functional teams including external stakeholders.
  • Demonstrated experience or potential for leadership, especially in making architectural decisions, coordinating the efforts of other engineers, and developing standards and best practices.

Nice To Haves

  • Currently hold or eligible to hold U.S. Security Clearance (Secret or higher)
  • Prior experience designing test and evaluation systems for AI
  • Experience deploying complex software systems in IL 4 or higher environments.
  • Experience participating in security assessment and authorization, hardening systems, and remediating vulnerabilities.
  • Previous experience contributing to open source projects and participating in open source communities.
  • Exposure to the intelligence community or department of defence subject matter experts.

Responsibilities

  • Design and implement a platform for test & evaluation of AI models and Agentic systems
  • Develop evaluation methodologies combining human and AI expert judging and multiple input and output modalities
  • Design and enhance user interfaces such that platform customers can express workflows in an intuitive and generic manner
  • Collaborate closely with subject matter experts, Model/Agent developers, and evaluation designers to ensure coverage over existing and to be discovered requirements.
  • Architect and implement robust experiment provenance and result tracking systems.
  • Ensure performance across integrated tools and scaling systems.
  • Provide technical mentorship, code review, design guidance, and clear documentation to support team knowledge transfer and user onboarding.

Benefits

  • 100% employer paid medical premiums for employees
  • self-managed PTO with a minimum time off requirement
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service