HPC Platform Product Lead

ZoetisKalamazoo, MI
Remote

About The Position

Owns the end-to-end success of Zoetis’ VMRD High Performance Computing (HPC) environment as both a product and a platform: sets vision and roadmap, defines service offerings and governance, and provides technical leadership for architecture, SLURM scheduling, reliability, and sustainable growth. Partners with scientific/analytics users and business leaders to translate workload demand into a prioritized backlog and measurable value narrative. Ensures the environment delivers the right performance, cost, and risk posture through capacity planning, operational excellence, and data lifecycle strategy.

Requirements

  • Bachelor’s degree in computer science, Engineering, Information Systems, or equivalent practical experience.
  • 7+ years of experience in infrastructure engineering/architecture, platform engineering, or adjacent technical product roles.
  • 5+ years of experience in HPC environments (on-prem, cloud, and/or hybrid).
  • Demonstrated experience partnering with scientists/engineers to translate workload needs into prioritized platform capabilities and measurable value stories.
  • Strong SLURM administration skills: partitions/queues, QoS, fairshare, accounting, reservations; job monitoring and troubleshooting; job efficiency analysis; user guidance.
  • HPC infrastructure knowledge: Linux administration fundamentals, compute (CPU/GPU), high-speed networking, shared/parallel storage concepts, and capacity/performance planning.
  • Automation/scripting (e.g., Bash/Python) and operational tooling; infrastructure-as-code/automation tooling where applicable.
  • Observability: utilization telemetry, job analytics, dashboards/alerting for node health and storage utilization; actionable operational reporting.
  • Data lifecycle and protection: retention policies, tiered storage, archive/backup strategy, restore testing, and operational runbooks.
  • Product and delivery skills: requirements discovery, value storytelling, roadmap/backlog prioritization, stakeholder management, and communication of technical trade-offs and ROI.

Nice To Haves

  • Master’s degree in a relevant field.
  • Experience operating HPC in regulated or highly governed environments.
  • Product Owner / Agile delivery experience (e.g., backlog management, acceptance criteria, road mapping).

Responsibilities

  • Own product vision, roadmap, service catalog, and backlog with clear acceptance criteria.
  • Translate utilization/throughput/outcomes into investment asks, funding recommendations, and executive-ready value stories (showback/chargeback as applicable).
  • Define target-state architecture and standards across compute, storage, network, security, and tooling; drive multi-year evolution planning.
  • Own scheduling strategy (partitions/queues, QoS, fairshare, accounting, reservations); lead configuration, troubleshooting, and job efficiency analysis; provide technical backup as needed.
  • Drive operational excellence via monitoring/alerting, dashboards, SLO/SLI practices, incident/problem follow-up, runbooks, and continuous improvement.
  • Plan and execute compute and storage growth; manage allocations; coordinate procurement and lifecycle refresh.
  • Establish onboarding/entitlements/RBAC and auditability; improve user support processes (ticketing, escalations, knowledge base) and publish workload best practices (CPU/GPU optimization, execution standards, data movement).
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service