Senior Product Manager, AI Infrastructure

Lightning AI•San Francisco, NY
•Hybrid

About The Position

Lightning AI is seeking a Senior Product Manager, AI Infrastructure to lead the product experience for their GPU cloud. This role involves defining how customers discover, provision, configure, and operate GPU compute, working closely with engineering and operations to enhance platform reliability, efficiency, and scalability. The ideal candidate understands that infrastructure is a product and is comfortable with concepts like GPU availability, cluster provisioning, networking, storage, scheduling, workload reliability, and customer-facing APIs and workflows. This position requires collaboration across Engineering, Infrastructure, Sales, customers, and the executive team to shape product strategy and differentiation. It's a high-ownership role involving direct problem investigation, data analysis for reliability and utilization, technical trade-offs with engineers, and driving products from conception through launch and iteration. The role reports to the VP of Product.

Requirements

  • 7+ years of product management experience, including 3+ years building cloud infrastructure, compute, platform, developer tooling, or AI infrastructure products.
  • Experience building technical products for developers, infrastructure teams, ML engineers, AI researchers, or other technical users.
  • Strong understanding of cloud infrastructure concepts including compute, networking, storage, provisioning, scheduling, and orchestration.
  • Familiarity with GPU infrastructure and the requirements of large-scale AI training, experimentation, or inference workloads.
  • Technical depth to work directly with engineers on APIs, distributed systems, Kubernetes, workload orchestration, observability, and infrastructure reliability.
  • Experience using data to understand infrastructure utilization, capacity, reliability, performance, and customer behavior.
  • Track record of owning technical products from problem definition through launch and adoption.
  • Strong product judgment and ability to turn complex infrastructure capabilities into simple customer experiences.
  • Experience with pricing, packaging, consumption-based products, or cloud infrastructure unit economics.
  • Strong prioritization skills and comfort making tradeoffs across customer needs, reliability, infrastructure efficiency, and engineering investment.
  • Strong written and verbal communication across technical, customer, and executive audiences.
  • Comfortable moving quickly and operating in ambiguous environments.
  • BS in Computer Science, Engineering, or equivalent practical experience.

Nice To Haves

  • Experience at a GPU cloud, neocloud, hyperscaler, AI infrastructure company, or infrastructure developer-tools company.
  • Experience building products involving GPU provisioning, cluster management, capacity management, workload scheduling, or distributed compute.
  • Familiarity with GPUs, Kubernetes, Slurm, Ray, PyTorch, distributed training, or similar infrastructure technologies.
  • Experience with reserved capacity, on-demand compute, utilization optimization, or other cloud consumption models.
  • Experience building infrastructure products that support large-scale AI training and inference.
  • Experience working closely with data center, hardware, networking, or infrastructure operations teams.

Responsibilities

  • Own the product vision and roadmap for Lightning AI’s GPU cloud infrastructure.
  • Define how customers discover, provision, configure, and consume GPU compute.
  • Build product experiences around GPU capacity, clusters, scheduling, networking, storage, and workload execution.
  • Partner closely with infrastructure and platform engineering to improve availability, reliability, utilization, and performance.
  • Develop a deep understanding of customer workloads (experimentation, training, inference) and translate needs into infrastructure capabilities.
  • Use infrastructure and product data to identify capacity constraints, reliability issues, performance bottlenecks, and opportunities for customer experience improvement.
  • Define the APIs, interfaces, and developer workflows for customer interaction with infrastructure.
  • Make product trade-offs across customer experience, infrastructure efficiency, reliability, cost, and engineering complexity.
  • Own pricing, packaging, and consumption models in partnership with Ops, Sales, and Finance.
  • Partner with GTM on positioning, technical sales conversations, customer feedback, and competitive differentiation.
  • Define and track metrics across GPU utilization, provisioning, workload reliability, infrastructure consumption, adoption, and retention.
  • Take products from problem discovery through requirements, launch, adoption, and iteration.

Benefits

  • Discretionary bonus
  • Meaningful equity component (RSUs)
  • Comprehensive Health Coverage: Medical, dental, and vision coverage for employees and eligible dependents.
  • Retirement Savings: 401(k) matching (U.S.) and pension contributions (U.K.).
  • Flexible Time Off: Unlimited PTO, company holidays, and floating holidays.
  • Company-Wide Winter Break: Two weeks of company closure each winter.
  • Paid Parental & Family Leave
  • Professional Development: Annual learning and development allowance.
  • Wellness Benefits: Wellness and work-from-home stipends.
  • Sabbatical Program: Four weeks of paid sabbatical leave after four years of service.
  • Flexible Work: Flexible schedules and a hybrid work model.
  • In-Office Meals: Complimentary meals at office hubs.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service