Corporate Vice President - Model Validation and AI Governance

New York LifeNew York, NY
$147,500 - $211,000Hybrid

About The Position

The Corporate Vice President – Model Validation and AI Governance will play a key leadership role in strengthening New York Life’s approach to model validation and responsible AI governance across predictive, generative, and agentic AI solutions. Working closely with Model Risk Management and partners across Artificial Intelligence & Data, Technology, Risk, Legal, Compliance, Cybersecurity, and Third-Party Risk Management, this role will translate model risk requirements into rigorous, practical validation approaches and controls. This role will lead complex validation engagements end-to-end for internally developed and third-party solutions, independently challenging model methodologies, evaluation approaches, controls, and monitoring strategies. A particular focus will be establishing robust approaches for evaluating agentic and multi-step AI systems, including tool use, orchestration, autonomy, guardrails, human oversight, observability, and emerging failure modes. The successful candidate will combine deep quantitative and AI expertise with strong risk judgment and communication skills. They will establish validation standards and reusable practices, mentor junior colleagues, and communicate technical findings, limitations, and conditions of use in clear terms that enable senior stakeholders and governance forums to make informed decisions.

Requirements

  • Advanced degree in Statistics, Computer Science, Data Science, Mathematics, Economics, Engineering, or a related quantitative discipline, with strong knowledge of statistics and 7+ years of experience in model validation, model governance, or model risk management for predictive and AI/ML solutions within regulated environments.
  • Deep hands-on knowledge of traditional statistical and machine learning approaches, combined with demonstrated experience validating generative and agentic AI systems, including multi-step workflows, tool use, orchestration and planning, guardrails, autonomy controls, human oversight, and agent-specific failure modes.
  • Demonstrated experience designing, building, or independently reviewing AI evaluation frameworks, including dataset curation, benchmark design, rubric-based and LLM-as-a-judge evaluation, human review, regression testing, retrieval-quality assessment, and adversarial or red-team testing.
  • Practical experience with AI/ML monitoring and observability, including tracing, logging, telemetry, dashboards, and alerting, as well as experience validating or overseeing third-party, vendor-hosted, foundation-model, or other limited-transparency AI solutions.
  • Proficiency in Python and SQL, with experience using agent orchestration, evaluation, observability, or related AI tooling and the technical depth to independently replicate results, conduct diagnostic analysis, and build challenger approaches when appropriate.
  • Strong communication, leadership, and analytical skills, including the ability to interpret technical and regulatory requirements, translate them into practical controls, lead complex validations end-to-end, mentor colleagues, and present and defend technical findings with senior leaders and cross-functional governance bodies.

Nice To Haves

  • Experience evaluating Generative AI solutions, including prompting strategies, retrieval-augmented generation (RAG), retrieval quality, model adaptation or fine-tuning trade-offs, and AI-specific privacy, security, interpretability, and stress-testing considerations.
  • Knowledge of agentic AI risks and controls, including prompt injection, excessive agency, tool and permission scoping, action reversibility, guardrail effectiveness, and hands-on exposure to commercial or open-source AI evaluation, observability, or guardrail technologies.
  • Familiarity with model and AI risk frameworks and emerging regulatory expectations, including SR 11-7, the NIST AI Risk Management Framework, the NAIC Model Bulletin on AI, and the EU AI Act, as well as experience collaborating with Legal, Compliance, Cybersecurity, and Third-Party Risk Management functions.
  • Applied understanding of AI use cases within financial services or insurance, such as underwriting support, fraud detection, marketing and sales enablement, agent productivity, and customer service, including the strengths, limitations, and evolving failure modes associated with generative and agentic AI.

Responsibilities

  • Lead independent validation and effective challenge across predictive, machine learning, generative AI, and agentic AI solutions, assessing model design, data and feature pipelines, methodologies, evaluation metrics, assumptions, limitations, and business impact. For agentic systems, evaluate architecture, planning and orchestration, tool use and permissions, retrieval and prompt design, memory and state, autonomy boundaries, guardrails, human-in-the-loop controls, and escalation and fallback mechanisms.
  • Design and advance rigorous AI evaluation practices by developing and challenging evaluation frameworks, benchmark and golden datasets, rubric-based and LLM-as-a-judge scoring, human review protocols, regression suites, offline and online testing, and adversarial or red-team evaluations. Assess statistical rigor, coverage, reproducibility, and the strength of validation evidence for high-impact use cases.
  • Establish monitoring, observability, and third-party validation approaches that address model drift, bias and fairness, stability, hallucinations, business outcomes, and agentic failure modes. Evaluate vendor and foundation-model solutions through independent testing and due diligence, and partner with teams to establish tracing, logging, telemetry, dashboards, alerts, compensating controls, and audit-ready evidence.
  • Translate risk and regulatory expectations into practical controls by interpreting technical standards, regulatory guidance, and internal procedures and converting requirements into validation methodologies, checklists, playbooks, standard operating procedures, monitoring expectations, and evidence requirements. Present validation findings, limitations, and conditions of use to senior stakeholders and governance committees.
  • Set standards and strengthen validation capabilities across teams by developing reusable templates, guidelines, test harnesses, and best practices for predictive, generative, and agentic AI; mentoring and reviewing the work of junior colleagues; partnering across data science, engineering, product, and control functions; and remaining current on evolving modeling techniques, AI research, evaluation methodologies, observability technologies, and governance practices.

Benefits

  • leave programs
  • adoption assistance
  • student loan repayment programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service