AI/ML Engineers at AI Validation & Monitoring (AVM) apply data, systems, computer science, clinical workflow, and governance expertise to help ensure that clinical AI products are supported by proportionate, traceable, and technically sound validation evidence. They work with AIA Governance Operations, clinicians, product teams, human-factors and user-experience specialists, data scientists, IT, architecture, patient safety, legal and regulatory partners, vendors, and other stakeholders to translate AIA Governance Policy and approved validation methods into clear evidence expectations and review conclusions. As the AI/ML Engineer - Validation & Evaluation, with the functional assignment of Validation Guidance, Guardrails, and Test Evidence, you will perform routine and moderately complex subject-matter review of the Validation, Performance, and Safety content and supporting evidence prepared. You will apply approved rubrics, evidence-sufficiency criteria, checklists, templates, and standard methods; identify missing evidence, unsupported conclusions, methodological limitations, and inconsistencies; document corrections, clarification questions, limitations, and escalation triggers; and draft traceable AVM concurrence or consultation comments. High-risk, novel, disputed, exceptional, or out-of-method questions are escalated to the Senior or Principal AI/ML Engineer comparisons, residual risk, and whether assessment conclusions are supported by the available evidence. Reviewing validation strategy and pathway selection, including retrospective analysis, prospective testing, simulation, user acceptance testing, human-factors evaluation, pilots, and other risk-proportionate methods, for alignment with intended use, workflow, and patient-safety risk. Evaluating performance expectations, metrics, thresholds, acceptance criteria, standard-of-practice Reviewing test plans and validation datasets for representativeness, traceability, limitations, production parity, functionality, robustness, calibration, uncertainty, bias, subgroup and equity evidence, and the adequacy of retesting evidence. Reviewing risk and safety controls, guardrails, task boundaries, safe refusal, escalation, fallback, human oversight, training, and other mitigations needed for the intended clinical or operational use. Evaluating human-computer interaction, clinical workflow, usability, automation-bias risk, patient-facing behavior, accessibility, and human-use evidence, and identifying when additional evaluation, clarification, or remediation is required. Assessing whether prior validation remains applicable after changes to the model, data, thresholds, workflow, population, interface, autonomy, guardrails, training, or intended use, and documenting needs for targeted retesting, re-evaluation, or revalidation. Producing clear, traceable review findings and recommended outcomes, including concurrence, concurrence with edits, revision required, additional evidence, consultation, escalation, or insufficient basis, and communicating technical findings to Product Leads and other technical and non-technical stakeholders. Maintaining and applying test-plan, validation-dataset, guardrail, subgroup, and production-parity templates; policy-to-evidence crosswalks; evidence examples; standard finding and clarification language; and recurring-gap themes that improve guidance, training, tools, and Product Lead self-sufficiency. This vacancy is not eligible for sponsorship/ we will not sponsor or transfer visas for this position. Also, Mayo Clinic DOES NOT participate in the F-1 STEM OPT extension program.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior