Expert Reliability Engineer

SAPReston, VA
$176,300 - $374,200Hybrid

About The Position

We are looking for an Expert Reliability Engineer (SRE) within the Shared Management Services (SMS) group in the Technology and Engineering unit of SAP Sovereign Cloud organization. In this role, you will join the Technology and Engineering team as a Site Reliability Engineer focused on securing and scaling the foundational platform that underpins SAP Sovereign Cloud. You will work alongside a globally distributed team of highly motivated engineers responsible for the design, development, deployment, and lifecycle management of the Sovereign Cloud Shared Management Services (SMS) platform, with a mandate that spans both operational reliability and platform security. As an expert Site Reliability Engineer, you will shoulder a shared ownership of the reliability, security, and operational excellence of production and non-production environments within Sovereign Cloud. It is expected that you will have extensive experience in all areas of a critical application administration stack spanning: source control (git), CI/CD platforms, identity and access management, secrets management, container orchestration, network security infrastructure, full-stack observability tooling across multi-cloud environments, and AI-assisted engineering workflows. You will lead an experienced team of globally distributed engineers in setting standards and best practices across responsibility areas including: Code reviews, Agentic coding (AI) security best practices, Agentic coding workflows and skill development, AI-assisted incident response, Unit test coverage, Functional test coverage, etc. You will treat security as a first-class reliability concern: hardening identity and access management, secrets management, and supply chain integrity are as central to this role as uptime and incident response. You will identify and close mission-critical capability gaps, define disciplined and standardized operational processes, and help the team navigate trade-offs across deployment plans, infrastructure investments, and day-to-day operational decisions. You will own and continuously improve backup and disaster recovery drills, ensuring the platform is failure-ready at global scale. You will work closely with Sovereign Cloud operations and engineering teams, regional counterparts, and hyperscaler provider partners.

Requirements

  • Ability to manage ambiguities while being innovative and collaborative
  • Extensive technology skills and the willingness to learn new topics quickly
  • Problem-solving, presentation, communication, and interpersonal skills
  • Ability to think strategically, delivering projects and work cross-organizationally
  • Knowledge of SAP and the SAP solution portfolio
  • Cultural awareness, intercultural competencies, and the ability to influence without formal authority
  • Ability to build trusted relationships with key stakeholders
  • Persistence, self-motivation, and willingness to work under pressure
  • Proven ability to work in cross-functional teams
  • Ability to lead and mentor junior engineers in setting and maintaining DevOps and SRE best practices
  • English (fluent)
  • 10+ years of experience in DevOps and/or SRE engineering
  • 7+ years of experience with a successful track record of leading engineering projects and cross-functional program teams
  • Deep mastery of SRE principles as they apply to a globally distributed, mission-critical platforms
  • Advanced Experience in architecture, engineering, and deployment of modern monitoring tooling such as Grafana, Promethius
  • Background in security or security-adjacent roles with a proven record of strong security foundational skills
  • Proven track record in end-to-end implementation initiatives
  • Experience in people management or staff level technical leadership
  • Expert level knowledge of Linux/Unix administration, networking fundamentals, and operating large-scale distributed cloud environments
  • Advanced proficiency with IaC, scripting, and automation such as Ansible, Terraform, Bash, Python, Powershell, CI/CD
  • Deep expertise with modern access control, authentication standards and identity federation at an enterprise scale.
  • Proven experience architecting AI-driven incident response workflows that reduce mean time to detection and resolution through automated anomaly correlation, intelligent alert triage, and LLM-assisted runbook execution
  • Proven ability to own the design, implementation, and continuous refinement of SRE reliability frameworks, including SLAs, SLOs, and SLIs, ensuring alignment between platform performance, business commitments, and engineering priorities
  • Willingness to subject yourself to a governmental security clearance process

Nice To Haves

  • AI-assisted engineering workflows
  • Agentic coding (AI) security best practices
  • Agentic coding workflows and skill development
  • AI-assisted incident response

Responsibilities

  • Shoulder a shared ownership of the reliability, security, and operational excellence of production and non-production environments within Sovereign Cloud.
  • Lead an experienced team of globally distributed engineers in setting standards and best practices across responsibility areas.
  • Treat security as a first-class reliability concern: hardening identity and access management, secrets management, and supply chain integrity.
  • Identify and close mission-critical capability gaps, define disciplined and standardized operational processes, and help the team navigate trade-offs.
  • Own and continuously improve backup and disaster recovery drills, ensuring the platform is failure-ready at global scale.
  • Work closely with Sovereign Cloud operations and engineering teams, regional counterparts, and hyperscaler provider partners.

Benefits

  • Constant learning, skill growth, great benefits, and a team that wants you to grow and succeed.
  • SAP North America Benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service