Principal AI Platform Operations Engineer

Western Governors UniversityRaleigh, NC
Remote

About The Position

The Principal AI Platform Operations Engineer is the highest technical individual contributor and authority within the AI Platform & Operations team. The Principal Engineer is responsible for setting the long-term AI operations and platform vision and strategy for WGU. This role operates at the intersection of frontier AI operations practices and enterprise-scale production systems, translating emerging developments into strategic technical investments that advance the organization's mission. The Principal AI Platform Operations Engineer shapes the AI operations discipline itself, defining the practices, standards, and culture that govern how the organization deploys and operates AI, while advising executive leadership, influencing the product roadmap, and representing WGU's AI platform capabilities externally.

Requirements

  • Recognized expertise across the full spectrum of AI operations: deployment, platform engineering, observability, reliability engineering, governance, and LLMOps/MLOps at enterprise scale.
  • Ability to define and communicate multi-year technical strategy; translate organizational goals into architectural vision and operational roadmaps.
  • Deep expertise in frontier AI operations practices and the ability to assess, adapt, and productionize emerging techniques at organizational scale.
  • Expert-level knowledge of AI safety, responsible AI, model risk, security, and enterprise AI governance including regulatory and compliance considerations.
  • Proven ability to influence organizational direction at the executive level; able to build consensus across competing priorities and stakeholders.
  • Strong platform and systems thinking: ability to design AI infrastructure that enables teams to move faster and build more reliably.
  • Track record of establishing engineering culture, standards, and practices that durably improve organizational capability.
  • Exceptional written and verbal communication; able to author technical strategy documents, present to boards and executives, and represent the organization externally.
  • Experience in technical hiring, team-building, and developing talent across multiple seniority levels.
  • Deep familiarity with the AI operations vendor landscape, open-source ecosystem, and technology frontier; able to make informed, timely, well-reasoned bets.
  • Strong cross-functional leadership; comfortable driving alignment across engineering, product, data science, legal, and executive stakeholders.
  • Master's Degree in Computer Science, Software Engineering, Data Science, Machine Learning, Math, Physics, or a related field
  • 7+ years of hands-on experience deploying and operating ML/AI systems in production at scale.
  • 5 years of experience working in an AI/ML context alongside Data Scientists or ML Engineers
  • 3 years of experience in building large-scale machine learning or deep learning models on a cloud platform
  • Demonstrated experience setting operational strategy or architectural vision at an organizational level.
  • Proven track record of delivering transformative AI platform or operations programs that produced measurable business impact.
  • Experience advising or influencing senior/executive leadership on AI operations strategy and investment.
  • Experience leading or significantly contributing to AI governance, responsible AI, or model risk frameworks.
  • Experience building and scaling AI engineering or operations teams, including hiring, competency-building, and culture-setting.
  • Experience representing an organization externally through conference talks, publications, or strategic partnerships.
  • Deep hands-on expertise in production AI reliability engineering, observability, and deployment at scale.

Nice To Haves

  • 10+ years of experience in software engineering, data science, or machine learning.
  • Experience with the Databricks platform; Databricks certifications (e.g., Databricks Certified Machine Learning Professional, Databricks Certified Data Engineer, or Databricks Certified Associate Developer for Apache Spark) are a plus.
  • AWS cloud platform experience; AWS certifications (e.g., AWS Certified Machine Learning, Specialty, AWS Certified DevOps Engineer) a plus.
  • Experience with infrastructure-as-code (Terraform, CloudFormation) and container orchestration (Kubernetes) for AI/ML workloads.
  • Experience in EdTech, personalized learning, or student-facing AI/ML platforms.
  • Experience with enterprise AI governance and compliance frameworks (FERPA, GDPR, etc.).
  • Contributions to open-source MLOps/LLMOps tooling, or published work and conference presentations in the AI/ML operations domain.
  • PhD in Computer Science, AI/ML, or a related field.
  • Experience operating in highly regulated environments (e.g., EdTech, healthcare, finance) with associated compliance requirements.
  • Recognized thought leadership in the AI/MLOps community through writing, speaking, or advisory roles.

Responsibilities

  • Defines the multi-year AI operations and platform technical vision and strategy for WGU, aligned to organizational mission and emerging technological opportunity.
  • Establishes the foundational reference architectures and engineering principles that govern how AI/ML systems are deployed, operated, and governed across the organization.
  • Leads the evaluation and adoption of frontier AI operations paradigms, tooling, and infrastructure; authors technical strategy documents and architectural decision records that guide org-wide direction.
  • Drives the most complex, highest-stakes AI operations programs in the organization, including large-scale platform, migration, and reliability initiatives.
  • Advises executives such as the Chief Technology Officer, VP of Engineering or Security, and senior product leadership on AI operations strategy, risk, and investment priorities.
  • Defines AI operations governance including responsible AI standards, model risk management, security, cost accountability, and compliance frameworks for the organization.
  • Identifies and incubates emerging operational capabilities; sponsors proof-of-concept initiatives that create future organizational leverage.
  • Serves as the primary external technical spokesperson for WGU's AI operations work; represents the organization at industry conferences, in partnerships, and with strategic vendors.
  • Designs and evolves the team's operating model, hiring criteria, competency framework, and engineering culture.
  • Develops Staff, Senior, and II-level engineers through technical mentorship, sponsorship, and organizational knowledge-building programs.
  • Authors internal and external thought leadership: technical blogs, whitepapers, architectural guides, and research contributions.
  • Partners with legal, compliance, and privacy teams to ensure AI systems meet regulatory requirements and institutional risk standards.
  • Performs other job-related duties as assigned.
  • This job description includes a general representation of job requirements rather than a comprehensive inventory of all required responsibilities or work activities. The contents of this document or related job requirements may change at any time with or without notice.

Benefits

  • medical, dental, vision, telehealth and mental healthcare
  • health savings account and flexible spending account
  • basic and voluntary life insurance
  • disability coverage
  • accident, critical illness and hospital indemnity supplemental coverages
  • legal and identity theft coverage
  • retirement savings plan
  • wellbeing program
  • discounted WGU tuition
  • flexible paid time off for rest and relaxation with no need for accrual
  • flexible paid sick time with no need for accrual
  • 11 paid holidays
  • other paid leaves, including up to 12 weeks of parental leave
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service