Kudu Dynamics is seeking a Site Reliability Engineer to build and operate complex compute platforms that remain reliable under challenging conditions. This role involves working across infrastructure, Linux systems, networking, distributed storage, observability, security, and deployment automation, with a strong emphasis on creating reproducible, maintainable, resilient, and easy-to-operate systems. A significant focus will be on Nix and NixOS, valuing declarative systems, reproducible environments, infrastructure-as-code, and minimizing configuration drift. Responsibilities include building NixOS-based servers, improving deployment pipelines, developing Nix modules, debugging distributed systems, and ensuring predictable platform rebuilds. The environment may include high-performance compute, distributed Linux file systems, network design, security-in-depth, ML/AI infrastructure, high-bandwidth data processing, cloud deployment, and fielded systems. The engineer will own solutions spanning CNO, analytics, security, infrastructure, and field deployments, contributing to customer operational insights from a multi-domain data environment.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level