We’re seeking an inference systems engineer to build the distributed serving and high-performance networking layer behind our large language models. You will own the path from model server to GPU fabric: prefill/decode disaggregation, KV-cache transfer, multi-node execution, and the observability and benchmarks needed to make those systems reliable in production. This role sits in the Inference team within the company’s Model team and partners closely with Model, Infrastructure, and Application engineering.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed