The Systems Engineering team owns the Linux kernel and host software stack underneath one of the largest GPU fleets in the world. When something breaks at the Kubernetes layer — a pod stuck in a mystery state, a cgroup that won't reclaim memory, a scheduler decision that makes no sense — the root cause almost always lives below the abstraction, in the kernel. Our team is the last line of defense: we go down into memory management, scheduling, and namespaces to find out what's actually happening, fix it, and often upstream the fix. As a Senior Software Engineer on the Systems Engineering team, you will be the kernel expert who traces complex, ambiguous failures in Kubernetes, pods, and containers back to their root cause in the Linux kernel — and fixes them there. You'll live at the boundary between kernel internals (cgroups, namespaces, the scheduler, memory management) and the container/orchestration stack built on top of them (containerd, runc, the kubelet, the CRI). Day-to-day, you'll debug kernel crashes and panics, chase down cgroup and namespace edge cases that only show up at fleet scale, and work with hardware, platform, and product teams to turn "the pod is doing something impossible" into a root cause and a patch — upstreaming it when it belongs in the broader Linux and Kubernetes ecosystems.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed