Principal AI Product Engineer

Nscale•Houston, NY

About The Position

Nscale is seeking a Principal AI Engineer (Specialised) to lead the inference and post-training pillar of their AI systems engineering organization. This role involves defining the multi-year technical roadmap for model serving, evaluation, and post-training on Nscale’s GPU cloud, encompassing dedicated and serverless inference and bring-your-own-model deployments. The Principal Engineer will lead significant architectural programs and set engineering standards for a team of 20-50+ engineers. This position requires deep technical expertise in AI systems, with decisions directly impacting the cost, latency, and reliability of Nscale's token serving and post-training workloads. The role spans the full stack, from kernel efficiency on GPU systems to fleet-level KV cache and serving architecture, model quality evaluation, and the interplay between inference and training hardware. The engineer will also frame solutions for the organization and define customer-facing API contracts.

Requirements

  • 10–15 years of engineering experience, with a clear track record of pillar-level impact on production AI systems.
  • 4+ years of hands-on work with LLMs in inference, GPU performance, evals, or post-training and RL, in production or research.
  • Demonstrated ability to define multi-year technical strategy for complex, multi-team AI systems organizations.
  • World-class depth in production LLM inference, GPU performance, evals, and/or post-training and RL infrastructure, with strong working knowledge across the rest.
  • Demonstrated ownership of the architecture of a large-scale production inference or training platform.
  • Proven ability to create architectural frameworks and engineering standards adopted across large engineering organizations.
  • Deep understanding of the hardware/software boundary for AI accelerators: CUDA or ROCm, memory bandwidth and interconnect constraints, and distributed compute paradigms.
  • Strong history of growing technical leaders (Staff and above) and multiplying technical capability across teams.
  • External recognition in the AI systems community through research, open source, or industry contribution.

Nice To Haves

  • Prior experience at a top-tier AI lab or major hyperscaler AI infrastructure team.
  • Maintainer or core contributor to a foundational inference, kernel, or RL framework (vLLM, SGLang, TensorRT-LLM, LMCache, FlashInfer, Triton, verl, OpenRLHF, TRL, DeepSpeed, Megatron-LM, etc.).
  • Hands-on depth in RL for LLMs (DPO/GRPO-style methods, reward modelling, multi-turn and tool-use RL) and the interaction between inference and training infrastructure.
  • Experience defining developer API platforms adopted at scale by external developers.
  • Deep experience with control plane / data plane architecture and cell-based deployment patterns in large-scale inference infrastructure.
  • Published work in AI systems: MLSys, NeurIPS Systems Track, OSDI, EuroSys, SC, or equivalent.
  • Experience with hardware-software co-design: custom accelerator kernels (CUDA, Triton), compiler-level optimization, AI hardware roadmap engagement.
  • Experience defining pricing, SLO, and capacity models for a commercial inference product.

Responsibilities

  • Define and own the multi-year technical roadmap for Nscale’s inference, evals, and post-training platform, and translate it into architecture that multiple teams can execute against.
  • Lead company-scale architectural initiatives in the pillar, such as next-generation serving (disaggregated prefill/decode, KV cache orchestration across GPU, host, and storage tiers, speculative decoding, multi-tenant scheduling), GPU kernel and model efficiency work (custom kernels, FP8/NVFP4/INT8/4 quantization, sparsity, distillation, MoE serving), evals and benchmarking frameworks, and post-training and RL infrastructure.
  • Establish engineering standards adopted across all AI teams: API design and compatibility guarantees, benchmarking and evals methodology, training stability norms, and performance testing practices.
  • Own the framework by which cost, latency, throughput, and model quality trade-offs are made and measured across the pillar.
  • Identify long-horizon systemic risks early (serving engine and framework bets, accelerator support, capability gaps) and resolve them before they block the organization.
  • Align AI engineering, research, product, and infrastructure leadership on multi-team technical strategy; frame technical trade-offs in product and commercial terms.
  • Mentor and develop Staff and Senior AI Engineers, and grow the next generation of inference technical leaders at Nscale.
  • Represent Nscale’s technical approach externally: open-source leadership in the frameworks we depend on, publications, conference talks, and partnerships with GPU vendors and AI labs.

Benefits

  • medical
  • dental
  • vision
  • flexible paid time off
  • parental leave
  • retirement plan participation
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service