Engineering Manager, Production Platform and Orchestration

Positron Corporation
•$200,000 - $300,000•Remote

About The Position

Positron is seeking an Engineering Manager to lead Production Platform and Orchestration within our Upstack Engineering organization. This team owns the software and operating practices that provision, deploy, observe, upgrade, and reliably operate Positron systems in production. You will inherit a technically strong core team and help it grow into a durable organization capable of supporting a fleet that is expanding by several multiples. This is a technical leadership role with real operational accountability. You will set direction, build the team, create clear ownership, and improve the systems and processes behind fleet orchestration, deployment lifecycle, observability, incident response, release automation, and production reliability. You will work closely with serving and API, model enablement, compiler and runtime, hardware, customer-facing, and data center partners. The strongest candidate will combine systems depth with organizational judgment, moving comfortably between architecture, delivery, incidents, people development, and cross-functional planning. This description intentionally emphasizes outcomes and ownership over a fixed organizational chart. As the fleet and customer base grow, the function may develop dedicated groups for fleet orchestration and capacity, deployment lifecycle, reliability and observability, data center operations, customer production operations, and operational tooling.

Requirements

  • Demonstrated success managing and growing engineering teams responsible for distributed systems, cloud infrastructure, production platforms, SRE, or a closely related domain.
  • Strong technical judgment across Linux systems, networking, orchestration, deployment systems, observability, and production reliability.
  • A proven record of turning ambiguous operational demands into a coherent roadmap, explicit ownership, and measurable engineering outcomes.
  • Ability to recruit, coach, and retain engineers across experience levels while maintaining a high technical bar.
  • Comfort operating across software, hardware, data center, customer, and business boundaries.
  • Excellent written and verbal communication skills, with sound prioritization and the ability to make tradeoffs visible to technical and executive stakeholders.
  • A hands-on leadership style, close enough to architecture and operations to ask the right questions without becoming the team's bottleneck.

Nice To Haves

  • Hands-on experience operating GPU, FPGA, ASIC, or other accelerator fleets in production.
  • Ownership of services with meaningful availability expectations, including on-call, incident management, root-cause analysis, and reliability planning.
  • Experience building control planes, schedulers, placement systems, capacity-management systems, or multi-rack orchestration.
  • Background in data center deployment, hardware lifecycle, firmware coordination, sparing and RMA processes, or production networking.
  • A track record of scaling infrastructure from early deployments to multiple sites or hundreds of systems.
  • Exposure to customer-facing infrastructure where engineering teams participate in production escalation and service readiness.
  • Demonstrated automation work that materially reduced operational toil, incident frequency, or recovery time.

Responsibilities

  • Lead, coach, and grow a team of engineers spanning production platform, fleet orchestration, reliability, and operational automation.
  • Hire thoughtfully, develop emerging leaders, and create ownership boundaries that remain effective as the organization scales.
  • Translate customer and business priorities into sequenced engineering work while protecting the team from reactive, unstructured operations.
  • Establish a clear technical and organizational roadmap for provisioning, deployment, environment lifecycle, fleet health, capacity, upgrades, and rollback.
  • Build reliable orchestration and control-plane capabilities for inventory, placement, configuration, health management, and multi-system operations.
  • Prepare the production platform for a heterogeneous accelerator environment that may include FPGA, ASIC, and GPU infrastructure.
  • Define service-level objectives, operational metrics, alerting standards, and a sustainable on-call model for customer-facing production systems.
  • Own the operating cadence for incidents, escalations, postmortems, corrective actions, launch readiness, and reliability reviews.
  • Increase automation across deployment, upgrades, remediation, diagnostics, capacity planning, and common support workflows so that fleet growth does not require linear headcount growth.
  • Partner with hardware and data center teams on rack bring-up, networking, firmware, sparing, failure handling, and platform transitions.
  • Collaborate with Distributed Serving and API, Model Enablement, and Compiler and Executor teams to turn new capabilities into supportable production services.

Benefits

  • Fully company-paid medical, dental, and vision insurance for you and your dependents
  • Company-paid life and disability coverage, with voluntary options to add more
  • Supplemental hospital, critical illness, and accident coverage available
  • Unlimited paid time off, we encourage everyone to truly unplug and recharge
  • 13 paid company holidays
  • Remote-first culture with a company-provided computer and home office setup
  • Competitive salary and equity
  • 401(k) with company matching, eligible from day one
  • Visa Support
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service