Senior Product Manager, Network

NscaleNew York, NY
$220,000 - $260,000

About The Position

Nscale is building a vertically integrated GenAI cloud platform, owning data centres, software, and applications for the AI stack using sustainable technology. The company culture emphasizes relentless innovation, ownership, accountability, excellence, urgency, openness, transparency, collaboration, adaptability, and resilience. Technical Product Managers at Nscale are responsible for the definition, delivery, and evolution of parts of the Nscale platform, working with engineering, design, research, and go-to-market teams. The Senior Technical Product Manager for Fleet Operations specifically owns the product strategy for the day 0–2+ operational software that manages the global GPU fleet, including bringing capacity online, maintaining its health, and restoring it quickly when issues arise. This involves partnering with Fleet Software engineering, SRE, and Support teams to develop solutions for provisioning, bringup, testing, deployment, monitoring, incident response, repair, RMA, firmware, and decommissioning. The role operates at a team scope, managing a major product area and driving multi-quarter initiatives to improve fleet availability, utilization, and time-to-recover.

Requirements

  • 5–8 years of product management experience in software or technology, with a track record of owning significant product areas in infrastructure, platform, or operations-facing products.
  • Strong technical fluency in large-scale systems: you can lead discussions with engineering on architecture, trade-offs, and feasibility across provisioning, orchestration, observability, and control-plane design
  • Experience building products for operators — SREs, NOC/support teams, data centre technicians, or similar — and a genuine appetite for understanding their workflows.
  • Demonstrated ability to move from an ambiguous operational problem space to shipped product outcomes that measurably improve reliability, efficiency, or time-to-recover.
  • Experience mentoring or informally leading peers.
  • Excellent written and verbal communication; you can make complex product decisions legible to engineers, operators, and executives alike.
  • Experience with data centre networking technologies, including high-performance GPU interconnects such as InfiniBand and RoCE (RDMA over Converged Ethernet), and an understanding of how backend (east–west/compute) and frontend (north–south/storage and management) network fabrics are designed and operated at scale.
  • Familiarity with WAN, edge, and global backbone architectures — including how multi-site connectivity, peering, and traffic engineering support a globally distributed GPU fleet.
  • Experience partnering with network engineering teams on fabric health, congestion monitoring, and link-level failure workflows, ideally in environments where network performance directly impacts training or inference workloads.

Nice To Haves

  • Degree in computer science, engineering, or a related field, or prior experience as an engineer or SRE.
  • Hands-on background in cloud infrastructure, bare-metal provisioning, fleet or hardware lifecycle management, observability/monitoring platforms, or incident management tooling.
  • Experience with bare-metal provisioning systems such as OpenStack Ironic (or equivalents like MAAS, Tinkerbell, or in-house provisioning stacks).
  • Experience with DCIM tools such as NetBox (or equivalents like Device42 or Nautobot) for inventory, cabling, and rack/asset management.
  • Experience with ITSM and ticketing platforms such as Jira Service Management (or equivalents like ServiceNow, Zendesk, or Freshservice) for support, incident, and RMA workflows.
  • Experience with observability and monitoring platforms such as Grafana, Prometheus, Datadog, or equivalents — ideally including defining SLOs, dashboards, and alerting for large fleets.
  • Familiarity with GPU or accelerated compute environments, data centre operations, or hyperscaler-style fleet management.
  • Experience operating in high-growth or early-stage environments where the product is being built alongside the fleet itself

Responsibilities

  • Own the strategy and roadmap for a significant Fleet Operations product area — e.g. provisioning and bring-up, fleet health and telemetry, incident and repair workflows, firmware and lifecycle management, or capacity and inventory.
  • Lead multi-sprint, cross-functional initiatives from problem framing through rollout across live GPU clusters, working hand-in-hand with Fleet Software, SRE, data centre operations, and Support.
  • Turn operational ambiguity into product: shadow on-call rotations, ride along with support and repair workflows, and translate recurring toil into tooling, automation, and platform capabilities.
  • Define the metrics that matter for a GPU fleet — availability, utilisation, MTTR, time-to-bring-up, hardware failure rates, support ticket deflection — and drive the roadmap against them.
  • Partner with engineering on architecture and trade-offs for systems that span bare metal, orchestration, observability, and control planes.
  • Drive incident reviews and postmortems into product commitments; close the loop so the same class of issue doesn't recur.
  • Mentor junior product managers and raise the quality bar for PRDs, reviews, and product decisions across the team.
  • Represent Fleet Operations in planning, reviews, and leadership updates.

Benefits

  • bonus
  • equity
  • commission programs
  • medical
  • dental
  • vision
  • flexible paid time off
  • parental leave
  • retirement plan participation
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service