Nscale is looking for a Staff AI Engineer (Specialized) to set technical direction for the inference and reinforcement learning systems at the core of our AI services platform — and for the APIs through which other engineers consume them. This role owns the architecture of how models are served on Nscale’s GPU cloud, how RL and post-training workloads run on it, and how both are exposed to customers and internal teams as clean, reliable, high-performance interfaces. You’ll work across 2–4 teams spanning serving, post-training, and platform, defining how these systems are built and establishing the standards that create engineering leverage across the organization. As a Staff engineer, you are the technical authority for this area of the AI stack. Your decisions determine the latency, throughput, and cost profile of every token Nscale serves, and the correctness and efficiency of every RL run on our platform. You resolve ambiguous architectural questions where the answer space is genuinely open — disaggregated versus co-located serving, on-policy versus off-policy RL infrastructure, where the API boundary should sit — and your solutions become the standards others build on.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed