Site Reliability Engineer

LeidosChantilly, VA
Hybrid

About The Position

Kudu Dynamics is seeking a Site Reliability Engineer to build and operate complex compute platforms that remain reliable under challenging conditions. This role involves working across infrastructure, Linux systems, networking, distributed storage, observability, security, and deployment automation, with a strong emphasis on creating reproducible, maintainable, resilient, and easy-to-operate systems. A significant focus will be on Nix and NixOS, valuing declarative systems, reproducible environments, infrastructure-as-code, and minimizing configuration drift. Responsibilities include building NixOS-based servers, improving deployment pipelines, developing Nix modules, debugging distributed systems, and ensuring predictable platform rebuilds. The environment may include high-performance compute, distributed Linux file systems, network design, security-in-depth, ML/AI infrastructure, high-bandwidth data processing, cloud deployment, and fielded systems. The engineer will own solutions spanning CNO, analytics, security, infrastructure, and field deployments, contributing to customer operational insights from a multi-domain data environment.

Requirements

  • Enthusiasm for learning new stuff.
  • Bachelor’s degree in Computer Science, Computer Engineering, a related field, or amazing equivalent skills.
  • Strong experience administering and troubleshooting Linux systems.
  • Experience with Linux networking, storage, and file systems.
  • Experience automating system configuration and deployment.
  • Experience with Nix and/or NixOS in production, lab, or significant personal environments.
  • Python, Go, Bash, or similar scripting/programming experience of 2+ years.
  • Ability to debug complex systems methodically across multiple layers of the stack.

Nice To Haves

  • Deep experience with NixOS, including custom modules, overlays, flakes, packaging, and reproducible deployments.
  • Experience operating fleets of Linux systems.
  • Infrastructure-as-code experience with Nix, Terraform, Ansible, or similar tooling.
  • Experience designing highly available or fault-tolerant systems.
  • Monitoring and observability experience with tools such as Prometheus, Grafana, Loki, OpenTelemetry, Elasticsearch, or similar systems.

Responsibilities

  • Own the reliability and operation of critical compute and data platforms.
  • Design, deploy, and maintain Linux infrastructure, including NixOS-based systems.
  • Build reproducible system configurations and deployment workflows using Nix and infrastructure-as-code.
  • Develop reusable NixOS modules, packages, flakes, and system configurations.
  • Improve platform resilience, fault tolerance, recoverability, and maintainability.
  • Build monitoring, logging, alerting, and observability capabilities that help us understand system behavior before customers notice problems.
  • Diagnose difficult failures across Linux, networking, storage, containers, and distributed systems.
  • Automate repetitive operational tasks and eliminate configuration drift.
  • Design and build security and maintainability components for deployed systems.
  • Help define operational standards, deployment practices, upgrade strategies, and disaster-recovery procedures.
  • Perform capacity planning and identify performance bottlenecks across compute, storage, and networking.
  • Own solutions spanning CNO, analytics, security, infrastructure, and field deployments.
  • Learn new skills and teach the rest of us what you know.

Benefits

  • Flexible work hours
  • Work-from-home options
  • Home-roasted coffee
  • Award-winning workspaces
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service