We are looking for an Expert Reliability Engineer (SRE) within the Shared Management Services (SMS) group in the Technology and Engineering unit of SAP Sovereign Cloud organization. In this role, you will join the Technology and Engineering team as a Site Reliability Engineer focused on securing and scaling the foundational platform that underpins SAP Sovereign Cloud. You will work alongside a globally distributed team of highly motivated engineers responsible for the design, development, deployment, and lifecycle management of the Sovereign Cloud Shared Management Services (SMS) platform, with a mandate that spans both operational reliability and platform security. As an expert Site Reliability Engineer, you will shoulder a shared ownership of the reliability, security, and operational excellence of production and non-production environments within Sovereign Cloud. It is expected that you will have extensive experience in all areas of a critical application administration stack spanning: source control (git), CI/CD platforms, identity and access management, secrets management, container orchestration, network security infrastructure, full-stack observability tooling across multi-cloud environments, and AI-assisted engineering workflows. You will lead an experienced team of globally distributed engineers in setting standards and best practices across responsibility areas including: Code reviews, Agentic coding (AI) security best practices, Agentic coding workflows and skill development, AI-assisted incident response, Unit test coverage, Functional test coverage, etc. You will treat security as a first-class reliability concern: hardening identity and access management, secrets management, and supply chain integrity are as central to this role as uptime and incident response. You will identify and close mission-critical capability gaps, define disciplined and standardized operational processes, and help the team navigate trade-offs across deployment plans, infrastructure investments, and day-to-day operational decisions. You will own and continuously improve backup and disaster recovery drills, ensuring the platform is failure-ready at global scale. You will work closely with Sovereign Cloud operations and engineering teams, regional counterparts, and hyperscaler provider partners.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed