About The Position

Reflection's Compute Platform team keeps our compute layer healthy and highly available. We run a Kubernetes-based platform distributed across multiple neo-clouds, where multi-cloud scheduling, node health, and performance debugging at scale present genuinely hard systems problems. As Compute Platform Lead, you'll provide front-line leadership of the team that builds and operates this layer. You'll build, mentor, and grow a team of strong systems engineers, guide the technical and architectural decisions across multi-cloud scheduling, cluster management, and next-generation GPU deployments, and work closely with our training teams to co-design fault tolerance, node health checks, and remediation. You'll stay close enough to the systems to make targeted contributions as an individual contributor and to maintain a deep understanding of the compute fleet our largest training runs depend on. Managing vendors — and the important deals that come with them — is a core part of the job.

Requirements

  • Experience building, mentoring, and growing systems or infrastructure teams while staying technically hands-on. (Comfortable growing into leading a team of ~10 quickly if you haven't managed at that scale before.)
  • Deep systems-level engineering experience with a focus on cluster-wide behavior and maintenance.
  • Strong coding ability and the credibility to earn the technical trust of a strong team.
  • Depth in at least one of orchestration, storage, or GPU hardware — with the ability to learn the rest. Deep GPU knowledge beyond standard Kubernetes (e.g., NCCL) is a plus, not a prerequisite.
  • Alignment with a Kubernetes-first architecture.
  • Cloud storage expertise — managing high-performance data products (like VAST) across multiple data centers and handling datasets and checkpointing at scale — is a plus.
  • Experience managing vendors, including negotiating and operating important deals.
  • Ability to guide strategy and drive execution across a multi-cloud, large-fleet environment, and to partner effectively with research and training teams.

Nice To Haves

  • Deep GPU knowledge beyond standard Kubernetes (e.g., NCCL) is a plus, not a prerequisite.
  • Cloud storage expertise — managing high-performance data products (like VAST) across multiple data centers and handling datasets and checkpointing at scale — is a plus.

Responsibilities

  • Build, mentor, and grow a high-performing team of systems engineers. Coach and support your reports in understanding, and pursuing, their professional growth.
  • Provide front-line leadership of engineering efforts to keep the compute fleet reliable and highly available — multi-cloud scheduling, cluster management, and the path to next-generation GPUs and increasingly larger cluster sizes.
  • Stay hands-on: become familiar with the team's technical stack enough to make targeted contributions as an individual contributor.
  • Manage day-to-day execution: prioritize the team's work and manage projects in a highly dynamic, fast-paced environment.
  • Guide technical and architectural decisions, emphasizing scalability, robustness, and reliability — automatic remediation, topology-aware scheduling, capacity planning, rapid hardware debugging, and cluster-wide monitoring and performance benchmarking.
  • Work closely with our training teams to co-design fault tolerance, node health checks, and remediation, and manage the vendor relationships and important deals the compute fleet depends on.
  • Prepare the fleet for what's next: next-generation GPUs and larger clusters, and — longer term — multi-cloud storage, petabyte-scale data replication, and GPU-to-GPU network performance.
  • Raise the bar for technical judgment, prioritization, communication, and execution in a fast-moving environment.

Benefits

  • Top-tier compensation: Salary and equity structured to recognize and retain our talent globally.
  • Stock options: Everyone who joins and contributes to Reflection's success gets to share in the upside through stock options.
  • Health & wellness: Comprehensive medical, dental, vision, and life, with an annual wellness allowance.
  • Meals: Lunch and dinner are provided in the office daily.
  • Life & family: 22 weeks paid parental leave for all new birthing and non-birthing parents, including adoptive and surrogate journeys.
  • Vacation days: Unlimited paid time off in the U.S. and 30 days in the U.K.
  • Sponsorship support: We sponsor visas to help exceptional talent join our team and support long-term immigration pathways where applicable.
  • Team building: We have regular off-sites, happy hours, and team celebrations.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service