Nscale is seeking a Principal AI Engineer (Specialised) to lead the inference and post-training pillar of their AI systems engineering organization. This role involves defining the multi-year technical roadmap for model serving, evaluation, and post-training on Nscale’s GPU cloud, encompassing dedicated and serverless inference and bring-your-own-model deployments. The Principal Engineer will lead significant architectural programs and set engineering standards for a team of 20-50+ engineers. This position requires deep technical expertise in AI systems, with decisions directly impacting the cost, latency, and reliability of Nscale's token serving and post-training workloads. The role spans the full stack, from kernel efficiency on GPU systems to fleet-level KV cache and serving architecture, model quality evaluation, and the interplay between inference and training hardware. The engineer will also frame solutions for the organization and define customer-facing API contracts.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Principal
Education Level
No Education Listed