Software Engineering Manager - Fleet Management

Nscaleβ€’Seattle, WA
β€’$300,000 - $350,000β€’Hybrid

About The Position

Nscale is seeking a Software Engineering Manager to lead the team responsible for building Fleet Manager, a workflow automation platform. This platform manages the provisioning, testing, and remediation of GPU nodes and network switches at scale. The role involves leading engineers who develop Python-based systems for the entire operational lifecycle of compute infrastructure, including device enrollment, burn-in testing, network configuration, GPU health monitoring, and automated remediation. The manager will own end-to-end delivery and team health, including planning, execution, hiring, and career development, while collaborating with Principal and Staff engineers on technical architecture. The culture is hands-on, with the manager expected to stay involved through design reviews, incident deep-dives, and code, with people leadership as the primary focus.

Requirements

  • 8+ years of software engineering experience building and operating production systems.
  • 2+ years directly managing software engineers.
  • Strong technical foundation in Python and distributed systems.
  • Credible in design and code reviews, and able to guide trade-offs in infrastructure automation or workflow tooling.
  • Track record of delivering complex projects from ambiguous requirements to production.
  • Hands-on day-2 operations experience (monitoring, incident response, performance optimization).
  • Proven people leadership: hiring, coaching, performance management, and developing engineers.
  • Driven by building distributed systems at scale, infrastructure reliability, scalability, security, and continuous improvement.
  • Experience using AI tools like Claude or Cursor as a core part of workflow and raising team leverage with them.
  • Ability to context-switch between technical depth, delivery judgment calls, and people leadership.
  • Excellent communication skills to build consensus with stakeholders.

Nice To Haves

  • Experience leading teams that build workflow orchestration systems (Temporal, Airflow, Prefect, or similar).
  • Hands-on experience with infrastructure tooling: DCIMs, NetBox, OpenStack, or ERP systems.
  • Bare-metal provisioning and automation: MAAS, Ironic, IPMI, PXE boot, or network automation.
  • Experience building or managing hardware lifecycle automation: provisioning, validation, testing, or remediation workflows.
  • GPU infrastructure experience: health monitoring, burn-in testing, or cluster management.
  • HPC and networking: datacenter topology, high-performance interconnects (InfiniBand, RoCE).
  • Working knowledge of Kubernetes, Infrastructure as Code (Terraform, Pulumi), AWS, and GCP.
  • Experience scaling a team through rapid growth: hiring pipelines, onboarding, and team processes that hold up under pressure.

Responsibilities

  • Lead and grow the team by hiring, onboarding, and developing software engineers.
  • Coach performance and career growth, set clear expectations, and manage performance.
  • Own on-call health, handover quality, and a sustainable operational load for the team.
  • Turn the Fleet Manager roadmap into executable plans, driving prioritization, managing dependencies and risks, and meeting commitments.
  • Run the team's planning, review, and escalation mechanisms, and remove blockers.
  • Communicate status, risks, and trade-offs clearly to stakeholders.
  • Uphold engineering standards, including code review, testing, CI/CD, incident response, and postmortems.
  • Own the operational health of the team's services, including SLOs, observability, alerting, and incident management.
  • Partner with Principal and Staff engineers on architecture and design decisions.
  • Stay hands-on by reviewing designs and code, digging into incidents, and writing code where it unblocks the team.
  • Collaborate with Product, Infrastructure, Platform, SRE, and UI/UX teams.
  • Champion AI tools like Claude and Cursor across the team.

Benefits

  • Highly competitive US compensation package (base + bonus + equity)
  • Performance reviews every 12 months
  • Dynamic progression plan tailored to ambitions
  • Flexible workplace
  • Medical insurance
  • Dental insurance
  • Vision insurance
  • Flexible paid time off
  • Parental leave
  • Retirement plan participation
Β© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service