AVP, AI/ML Ops

Lincoln Financial•Radnor, PA
•Hybrid

About The Position

We are creating a new AI/ML Ops pillar to own the reliability, safety, and cost-effectiveness of every model, pipeline, and agent we run in production. This is fundamentally an operations leadership role: someone who has built monitoring, escalation, and change management discipline from the ground up during a technology or business transformation, and who can apply that same rigor to AI/ML systems. This leader will build and run the function that closes the gap between "built" and "running well at scale." They will report to the Chief Data & AI Engineering Officer and sit alongside the leaders of AI/Data Infrastructure, Data Strategy & Governance, Data & Delivery Platform, AI/ML/Agentic Engineering, and AI Native Intelligence Platform.

Requirements

  • Proven track record leading operations through a major transformation: building the operating model, monitoring and observability foundation, and escalation discipline for a function that didn't have one before, ideally in technology, platform, or shared-services operations.
  • Deep experience in change management: standing up change advisory processes, risk tiering, release governance, and rollback authority in a fast-moving technical environment.
  • Strong background in monitoring, observability, and notification and escalation design: knowing what to alert on, who gets paged, and how severity and ownership are defined.
  • Experience building incident response and postmortem practices, including on-call rotations and MTTD/MTTR reduction, in any production technology environment. AI/ML-specific experience is a strong plus, not a prerequisite.
  • Working knowledge of (or demonstrated ability to quickly learn) ML/LLM system failure modes: drift, degradation, hallucination, latency and cost blowups, and how to instrument against them.
  • Experience partnering with governance and risk functions to enforce policy in live systems.
  • Strong track record hiring, developing, and leading technical or operational teams, ideally from an early stage.
  • Comfortable operating with ambiguity and building process from scratch, not just running an existing playbook.

Responsibilities

  • Define the operating model for how AI/ML systems move from build to production: intake, change approval, release governance.
  • Build change management processes for model, pipeline, and agent releases, including change advisory review, risk tiering, rollback authority, and communication protocols.
  • Establish a clear ownership and accountability model (RACI) across Engineering, Ops, and Governance for every production AI/ML asset.
  • Build monitoring for performance, drift, latency, quality, and safety across production AI/ML systems.
  • Design the notification and alerting architecture: what gets flagged, to whom, at what severity, with what SLA for acknowledgment.
  • Implement in-production evals and guardrails; enforce governance policy at runtime in partnership with Data Strategy & Governance.
  • Define and own SLOs/SLAs for production models, pipelines, and agents.
  • Build on-call rotations, incident response, escalation tiers, and postmortem practice for AI/ML systems.
  • Own the detection and escalation path for misuse, abuse, or safety incidents.
  • Drive mean-time-to-detect and mean-time-to-recovery down as the estate scales.
  • Partner with Platform Engineering to stand up CI/CD, versioning, and rollback for models, pipelines, and agents, partnering with Engineering on execution.
  • Establish safe rollout patterns (canary, shadow, staged) in partnership with Engineering.
  • Own the retirement and deprecation process for models and agents no longer in use.
  • Own inference and compute cost visibility and optimization; build unit economics per model and agent.
  • Partner with AI/Data Infrastructure on capacity planning.
  • Define the operating model, team structure, and hiring plan for the pillar.
  • Build the feedback loop that routes production failure modes and cost signals back to Engineering.
  • Establish the pillar's KPIs and report on them to the Chief Data & AI Engineering Officer and broader leadership.

Benefits

  • PTO/parental leave
  • Competitive 401K and employee benefits
  • Free financial counseling, health coaching and employee assistance program
  • Tuition assistance program
  • Work arrangements that work for you
  • Effective productivity/technology tools and training
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service