Kudu Dynamics is seeking a Site Reliability Engineer to build and maintain reliable systems in complex environments. This role involves working across infrastructure, Linux systems, networking, distributed storage, observability, security, and deployment automation, with a focus on creating reproducible, maintainable, resilient, and easy-to-operate systems. A significant part of the role will utilize Nix and NixOS, emphasizing declarative systems, reproducible environments, infrastructure-as-code, and reducing configuration drift. Responsibilities include building NixOS-based servers, improving deployment pipelines, developing Nix modules, debugging distributed systems, and ensuring predictable platform rebuilds. The environment may include high-performance compute, distributed Linux file systems, network design, security-in-depth, ML/AI infrastructure, high-bandwidth data processing, cloud deployment, and fielded systems. This is an opportunity to own solutions and learn adjacent areas while ensuring customers receive operational insights from a dependable, observable, reproducible, and recoverable infrastructure.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level