Senior Machine Learning Engineer, AI Evaluation

SHRMAlexandria, VA
Hybrid

About The Position

The Senior Machine Learning Engineer, AI Evaluation builds and operates the measurement and engineering infrastructure supporting the organization's Applied AI Research (AAIR) function. This role is responsible for designing and maintaining the engineering infrastructure used to conduct rigorous, reproducible AI model evaluations and benchmarks. The Senior Machine Learning Engineer develops the systems that run multiple AI models against structured, domain-specific evaluations; builds scoring and evaluation frameworks; maintains reproducibility across model versions; and creates the data infrastructure necessary to analyze and track model performance over time. Working closely with HR subject matter experts and Applied AI Research colleagues, this position translates expert-defined standards and evaluation criteria into technically rigorous, measurable specifications. The position serves as a shared technical engineering resource across multiple Applied AI Research workstreams and helps ensure that published findings, benchmarks, and research conclusions are supported by reliable, auditable, and defensible measurement practices. This is an AI evaluation and engineering infrastructure role rather than a model-training or frontier AI research position.

Requirements

  • Bachelor's degree in Computer Science, Data Science, Machine Learning, Engineering, or a related quantitative or technical field, or relevant equivalent experience in lieu of degree.
  • Seven (7) or more years of progressively responsible experience in ML/LLM engineering, applied AI, applied data science, or research infrastructure, including experience developing, implementing, and supporting production-grade systems.
  • Demonstrated hands-on experience developing multi-model LLM applications and infrastructure, including provider-agnostic model access, APIs, prompt engineering, and evaluation frameworks.
  • Demonstrated experience with AI model evaluation and benchmarking, including rubric-based scoring, inter-rater reliability, model-as-judge methodologies and their limitations, and approaches for evaluating performance when definitive ground truth may not be readily available.
  • Experience designing and maintaining reproducible technical systems incorporating version control, model-version pinning, comprehensive logging, experiment tracking, and/or drift detection.
  • Experience with cloud-based data and AI infrastructure on a major cloud platform; Google Cloud Platform experience, including BigQuery, Vertex AI, IAM, and audit logging, preferred.
  • Experience developing and supporting data pipelines, structured experiment repositories, dashboards, or monitoring solutions.
  • Advanced proficiency in Python and strong software-engineering fundamentals, including the ability to develop reliable, maintainable, production-quality code.
  • Strong knowledge of machine learning, large language models, generative AI systems, and contemporary AI application architectures.
  • Demonstrated knowledge of AI evaluation and benchmarking methodologies, including rubric-based evaluation, automated scoring, model-as-judge approaches, inter-rater reliability, and measurement design.
  • Strong understanding of the limitations and failure modes of generative AI systems and the ability to design evaluation approaches that appropriately account for those limitations.
  • Demonstrated commitment to reproducibility, including disciplined use of versioning, documentation, logging, experiment tracking, and drift detection.
  • Ability to translate complex, judgment-based requirements from subject matter experts into technically rigorous and measurable evaluation specifications without oversimplifying the underlying domain expertise.
  • Strong analytical and problem-solving skills with the ability to identify technical, methodological, and data-quality issues and develop appropriate solutions.
  • Working knowledge of cloud-based AI and data environments, APIs, data warehouses, access controls, and related technical infrastructure.
  • Ability to effectively communicate complex technical concepts, methodologies, limitations, and findings to technical and non-technical audiences.
  • Strong collaboration and consultation skills, with the ability to work effectively with researchers, engineers, data professionals, subject matter experts, and external partners.
  • Ability to balance technical rigor, research requirements, scalability, and practical implementation considerations.
  • Strong understanding of responsible AI principles, data governance, privacy, security, and appropriate handling of sensitive or proprietary information.
  • Ability to evaluate emerging AI models, technologies, and evaluation methodologies and determine their appropriate application within the organization's research environment.
  • Ability to effectively leverage AI tools and technologies to streamline workflows, enhance productivity, and improve overall work quality.
  • Prolonged periods of sitting at a desk and working on a computer.
  • Frequent use of hands and fingers for typing, handling documents, and using office equipment.
  • Occasional standing, walking, bending, and reaching.
  • Ability to lift and carry up to 30 pounds as needed.
  • Clear verbal and written communication skills for effective interaction with colleagues and stakeholders.

Nice To Haves

  • Master's degree in Computer Science, Data Science, Machine Learning, Artificial Intelligence, or a related field preferred.
  • Experience working with HR, workforce, survey, behavioral, or other professional-domain data preferred.
  • Experience supporting academic, applied research, benchmarking, or peer-reviewed research workflows preferred.

Responsibilities

  • Design, build, and maintain scalable engineering infrastructure for conducting structured evaluations and experiments across multiple AI and large language model (LLM) families.
  • Develop and maintain a unified, provider-agnostic orchestration layer that enables consistent evaluation across multiple frontier model providers and architectures.
  • Design and implement rigorous AI evaluation and scoring frameworks, including rubric-based scoring, model-as-judge methodologies with appropriate safeguards, partial-credit methodologies, and approaches for managing ambiguity.
  • Build systems and processes that support reproducible experimentation, including model-version pinning, comprehensive run logging, experiment tracking, and drift detection.
  • Maintain portable evaluation architecture across AI model providers to enable consistent and defensible cross-model comparisons as models and technologies evolve.
  • Establish and maintain technical standards and engineering practices that support reliable, repeatable, and auditable AI evaluation.
  • Partner closely with HR subject matter experts to translate professional standards, research criteria, and judgment-based rubrics into measurable and technically executable evaluation specifications.
  • Identify and surface ambiguity, inconsistencies, or measurement limitations within proposed evaluation criteria and collaborate with subject matter experts to strengthen evaluation design.
  • Apply knowledge of AI evaluation methodologies, benchmarking techniques, inter-rater reliability, and known limitations of automated and model-as-judge evaluation approaches.
  • Support the design of measurement methodologies when definitive ground truth is unavailable or requires expert interpretation.
  • Ensure evaluation methodologies align with research-defined validation standards and produce findings that are reproducible, transparent, and defensible.
  • Contribute technical expertise to the design and continuous improvement of AI research experiments, benchmarks, and evaluation methodologies.
  • Design and maintain structured repositories for experiment results, prompt libraries, scoring rubrics, model metadata, and longitudinal evaluation data using cloud-based data infrastructure.
  • Build and maintain data structures in BigQuery or comparable platforms that enable research results to be queried, analyzed, reproduced, and audited.
  • Develop monitoring, reporting, and visualization capabilities using Looker, Looker Studio, or comparable tools to provide visibility into experiment status, model performance, and performance drift.
  • Establish processes for tracking changes in model behavior across model versions and over time.
  • Maintain complete technical documentation and metadata necessary to reproduce research findings and evaluation results.
  • Ensure appropriate quality controls are incorporated throughout data collection, evaluation, scoring, storage, and reporting processes.
  • Serve as a shared engineering resource across multiple Applied AI Research teams and research workstreams.
  • Collaborate with research leaders, HR subject matter experts, data professionals, engineers, and other internal stakeholders to translate research requirements into scalable technical solutions.
  • Support collaboration with university, affiliate, research, and other external partners when appropriate and within established organizational access controls and data-handling requirements.
  • Communicate technical methodologies, limitations, risks, and findings clearly to both technical and non-technical audiences.
  • Evaluate emerging AI models, tools, technologies, and evaluation methodologies and recommend appropriate applications within the research environment.
  • Contribute to continuous improvement of the Applied AI Research technical environment, engineering practices, and evaluation capabilities.
  • Ensure AI evaluation systems and workflows comply with organizational requirements for data security, privacy, ownership, access, and responsible AI use.
  • Maintain appropriate controls to protect proprietary, member, research, and other sensitive data from unauthorized access or use.
  • Ensure organizational data is not used to train external shared models except where expressly authorized and appropriately governed.
  • Implement and maintain appropriate access controls and data-handling requirements when working with external research or partner organizations.
  • Partner with appropriate internal stakeholders to ensure evaluation infrastructure aligns with organizational technology, security, privacy, and governance standards.

Benefits

  • professional growth and development
  • health
  • dental
  • vision
  • well-being
  • health savings
  • flexible spending
  • retirement
  • open leave
  • annual discretionary bonus and incentives
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service