This role is for a Senior Software / Site Reliability Lead Engineer in the defense industry. The engineer will be responsible for setting and enforcing cross-pod reliability standards for AI services, defining and managing Service Level Objectives (SLOs) and error budgets, implementing and maintaining the full observability stack (logging, metrics, tracing, dashboards), designing and managing alerting infrastructure, owning incident response procedures, and defining and enforcing production readiness criteria for AI services. The role also involves identifying and automating repetitive operational tasks (toil elimination). A key differentiator of this role is the application of SRE principles from scratch for a new platform, with the engineer having direct authority over whether AI services go live. The position requires a strong software engineering background to collaborate with development teams at the design level and address reliability issues proactively. The engineer will focus on unique AI service failure modes such as model drift and token budget exhaustion, which are not typically encountered by traditional SRE teams.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior