About The Position

Infrastructure Reliability Engineering (IRE) is a small but growing team responsible for the infrastructure and operations behind the core developer tools and on-prem compute platforms used across the entire engineering organization. We own the services every engineer depends on daily — source control, CI/CD, and artifact management — as well as the on-prem infrastructure behind simulation and GPU workloads. As the company’s on-prem footprint grows, this team is expanding its scope to provide SRE capabilities for on-prem systems, so there’s an opportunity to help shape that practice from the ground up. You’ll own the full lifecycle — patching, upgrades, backups, scaling, and incident response — for services that engineering depends on daily. The role blends DevOps, SRE, and software engineering, and is ideal for engineers who want high ownership and company-wide impact. You should have a mindset of continuous improvement: if something is manual and repetitive, your instinct should be to automate it away.

Requirements

  • Experience operating infrastructure outside of managed cloud services — bare-metal kubernetes and on-prem virtualization (VMware ESXi/vSphere)
  • Experience operating production systems using Docker and Kubernetes
  • Strong foundational knowledge of Linux (RHEL , Ubuntu)
  • Proficiency with at least one cloud platform (AWS, GCP, or Azure)
  • Experience managing infrastructure with Infrastructure-as-Code tools (e.g., Terraform/OpenTofu)
  • Experience with configuration management tooling (e.g., Ansible, Puppet, Chef)
  • Strong problem-solving skills with a focus on automation
  • Scripting or software development experience (e.g., Python, Go, Bash)
  • Familiarity with CI/CD pipelines and developer tooling
  • Ability to own systems end-to-end, from design to incident resolution
  • Eligible to obtain and maintain an active U.S. Secret security clearance

Nice To Haves

  • Experience with RKE2 (or other bare-metal Kubernetes distro such as k3s, kubeadm, or OpenShift) and Cilium
  • Prior experience with GitHub Enterprise Server, JFrog Artifactory/Xray, or CircleCI
  • Experience with GitOps workflows and tooling (e.g., ArgoCD/FluxCD)
  • Experience maintaining highly available, scalable internal tools
  • Exposure to security best practices, compliance requirements, or auditing
  • Experience supporting large, rapidly scaling engineering organizations
  • Experience with monitoring and observability platforms (e.g., Datadog, Prometheus, Grafana)
  • Background in SRE or hybrid SWE/DevOps roles
  • Experience with on-prem infrastructure operations, reliability, or capacity planning
  • Experience operating GPU or HPC infrastructure and workload schedulers (eg., RunAI, Slurm, Kubeflow, Volcano)

Responsibilities

  • Serve as a primary owner for critical services, including on-call and knowledge-sharing across the team
  • Own the lifecycle of core self-hosted developer tools (e.g., RunAI, GitHub Enterprise Server, CircleCI, JFrog Artifactory/Xray)
  • Design and implement automated systems for patching, backups (with validation), and upgrades
  • Scale infrastructure to support a fast-growing engineering org
  • Use Infrastructure-as-Code (Terraform) to manage environments
  • Operate and troubleshoot systems using Docker, Kubernetes, and cloud platforms (AWS, GCP, Azure)
  • Define and maintain SLOs for service availability, reliability, and performance
  • Build and maintain monitoring, alerting, and observability for developer tool services
  • Lead and participate in incident response and root cause analysis
  • Work cross-functionally with platform, security, infrastructure (on-prem and cloud), and software teams

Benefits

  • Highly competitive equity grants are included in the majority of full time offers; and are considered part of Anduril's total compensation package.
  • top-tier benefits for full-time employees
  • comprehensive, competitive benefits package (available at little to no cost to employees) ensures you’re supported in health, recovery, and whatever comes next.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service