Senior Software / Site Reliability Lead Engineer

General Dynamics Mission Systems, Inc,
$142,696 - $158,303Remote

About The Position

This role is for a Senior Software / Site Reliability Lead Engineer in the defense industry. The engineer will be responsible for setting and enforcing cross-pod reliability standards for AI services, defining and managing Service Level Objectives (SLOs) and error budgets, implementing and maintaining the full observability stack (logging, metrics, tracing, dashboards), designing and managing alerting infrastructure, owning incident response procedures, and defining and enforcing production readiness criteria for AI services. The role also involves identifying and automating repetitive operational tasks (toil elimination). A key differentiator of this role is the application of SRE principles from scratch for a new platform, with the engineer having direct authority over whether AI services go live. The position requires a strong software engineering background to collaborate with development teams at the design level and address reliability issues proactively. The engineer will focus on unique AI service failure modes such as model drift and token budget exhaustion, which are not typically encountered by traditional SRE teams.

Requirements

  • Bachelor’s degree in Computer Science, Software Engineering, or a related field, plus 8 years of experience; or Master’s degree plus 6 years of experience
  • Production SRE or DevOps experience — you have owned the reliability of systems that real users depended on, not just built CI/CD pipelines
  • Hands-on experience with monitoring and observability tools — Prometheus, Grafana, Datadog, ELK, CloudWatch, or similar. You have built dashboards and alerts that caught real problems.
  • Strong scripting and automation skills — Python, Bash, infrastructure-as-code (Terraform, CloudFormation, or similar)
  • Experience with containerized environments — Docker, Kubernetes, container orchestration at scale
  • Experience defining and managing SLOs, error budgets, and incident response procedures in production
  • U.S. citizenship required.
  • Department of Defense Secret security clearance is required at time of hire.

Nice To Haves

  • Production SRE or DevOps experience — you have owned the reliability of systems that real users depended on, not just built CI/CD pipelines
  • Software engineering fundamentals — you can read, write, and meaningfully review production-quality code. You understand how architectural and design decisions made early translate into operational problems later.
  • Software design experience — you have participated in or led design reviews, defined service interfaces or APIs, and pushed back on design decisions using reliability and operability as criteria
  • Hands-on experience with monitoring and observability tools — Prometheus, Grafana, Datadog, ELK, CloudWatch, or similar. You have built dashboards and alerts that have caught real problems.
  • Strong scripting and automation skills — Python, Bash, infrastructure-as-code (Terraform, CloudFormation, or similar)
  • Experience with containerized environments — Docker, Kubernetes, container orchestration at scale
  • Experience defining and managing SLOs, error budgets, and incident response procedures in production
  • You build things that work. Your default response to a problem is code, not a document.
  • You have shipped AI systems that real users depended on in production.
  • You are comfortable working without detailed specs — you can take a problem statement and figure out the right approach.
  • You care about reliability as much as capability — you monitor what you deploy.
  • You move fast without being reckless. You know when to iterate and when to get it right the first time.

Responsibilities

  • Cross-pod reliability standards. Set the reliability bar and ensure it is met consistently across applications.
  • SLOs and reliability metrics. Own definitions of service level objectives for every AI service that goes to production. Establish error budgets and use them to drive engineering decisions — not just measure uptime.
  • Monitoring and observability. Implement and maintain the full observability stack — logging, metrics, tracing, and dashboards.
  • Design and manage alerting infrastructure that tells you what's wrong, not just that something is wrong.
  • Incident response. Own on-call procedures, escalation paths, and incident management end-to-end. Lead post-incident reviews and maintain the reliability improvement backlog.
  • Production Readiness. Define and enforce the criteria that determine whether an AI service is ready for production.
  • Toil elimination. Identify and automate repetitive operational tasks.

Benefits

  • highly competitive benefits
  • flexible work environment where contributions are recognized and rewarded
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service