Site Reliability Engineer

LeidosBoulder, CO
$87,100 - $157,450Hybrid

About The Position

Kudu Dynamics is seeking a Site Reliability Engineer to build and maintain reliable systems in complex environments. This role involves working across infrastructure, Linux systems, networking, distributed storage, observability, security, and deployment automation, with a focus on creating reproducible, maintainable, resilient, and easy-to-operate systems. A significant part of the role will utilize Nix and NixOS, emphasizing declarative systems, reproducible environments, infrastructure-as-code, and reducing configuration drift. Responsibilities include building NixOS-based servers, improving deployment pipelines, developing Nix modules, debugging distributed systems, and ensuring predictable platform rebuilds. The environment may include high-performance compute, distributed Linux file systems, network design, security-in-depth, ML/AI infrastructure, high-bandwidth data processing, cloud deployment, and fielded systems. This is an opportunity to own solutions and learn adjacent areas while ensuring customers receive operational insights from a dependable, observable, reproducible, and recoverable infrastructure.

Requirements

  • Enthusiasm for learning new stuff.
  • Bachelor’s degree in Computer Science, Computer Engineering, a related field, or amazing equivalent skills.
  • Strong experience administering and troubleshooting Linux systems.
  • Experience with Linux networking, storage, and file systems.
  • Experience automating system configuration and deployment.
  • Experience with Nix and/or NixOS in production, lab, or significant personal environments.
  • Python, Go, Bash, or similar scripting/programming experience of 2+ years.
  • Ability to debug complex systems methodically across multiple layers of the stack.

Nice To Haves

  • Deep experience with NixOS, including custom modules, overlays, flakes, packaging, and reproducible deployments.
  • Experience operating fleets of Linux systems.
  • Infrastructure-as-code experience with Nix, Terraform, Ansible, or similar tooling.
  • Experience designing highly available or fault-tolerant systems.
  • Monitoring and observability experience with tools such as Prometheus, Grafana, Loki, OpenTelemetry, Elasticsearch, or similar systems.

Responsibilities

  • Own the reliability and operation of critical compute and data platforms.
  • Design, deploy, and maintain Linux infrastructure, including NixOS-based systems.
  • Build reproducible system configurations and deployment workflows using Nix and infrastructure-as-code.
  • Develop reusable NixOS modules, packages, flakes, and system configurations.
  • Improve platform resilience, fault tolerance, recoverability, and maintainability.
  • Build monitoring, logging, alerting, and observability capabilities that help us understand system behavior before customers notice problems.
  • Diagnose difficult failures across Linux, networking, storage, containers, and distributed systems.
  • Automate repetitive operational tasks and eliminate configuration drift.
  • Design and build security and maintainability components for deployed systems.
  • Help define operational standards, deployment practices, upgrade strategies, and disaster-recovery procedures.
  • Perform capacity planning and identify performance bottlenecks across compute, storage, and networking.
  • Own solutions spanning CNO, analytics, security, infrastructure, and field deployments.
  • Learn new skills and teach the rest of us what you know.

Benefits

  • Flexible work hours
  • Work-from-home options
  • Home-roasted coffee
  • Award-winning workspaces
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service