Senior Site Reliability Engineer

OktaSan Francisco, WA
Remote

About The Position

Okta's Technology, Data and Intelligence (TDI) team delivers the systems, tools, and services that power internal operations across the company. From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology. The Senior Site Reliability Engineer Opportunity Reporting to the Manager, Site Reliability Engineering, this role will help build, improve, and maintain our cloud platform services by designing and implementing complex cloud-based engineering enablement systems. With a strong focus on automation, testing, and operational excellence, you will deliver foundational infrastructure capabilities that enable corporate engineering teams to operate securely, reliably, and at scale.

Requirements

  • 5+ years of experience in SRE, DevOps, or Systems Engineering roles with a proven track record of delivering complex, large-scale infrastructure projects.
  • Expert in building and managing AWS multi-account environments (spanning hundreds of accounts), with deep proficiency in authentication, governance, and organization management (AWS Orgs, IAM, Identity Center, StackSets).
  • Highly skilled in infrastructure as code (Terraform), writing secure automation tools in Python, and building Git-based CI/CD workflows (GitLab, GitHub Actions).
  • Strong hands-on experience managing container orchestration environments (Kubernetes) and utilizing monitoring and logging tools (Splunk, CloudWatch, Grafana stack).
  • This position requires the ability to access federal environments and/or have access to protected federal data. As a condition of employment for this position, the successful candidate must be able to submit documentation establishing U.S. Person status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee. 22 CFR 120.15) upon hire.

Nice To Haves

  • Hands-on experience with general networking concepts (BGP and IPsec management) and leveraging core AWS networking services (VPCs, TGWs, and VPC endpoints).
  • Solid foundational knowledge and experience in Linux system administration.
  • Proven experience operating within highly secure, regulated environments (e.g., FedRAMP), with a strong understanding of FIPS, STIGs, and data boundary implementations.

Responsibilities

  • Design, build, and modernize scalable cloud environments and development tools while strictly enforcing security policies and standards for regulated environments.
  • Partner with software engineering teams to champion DevOps and SRE best practices, deliver excellent internal customer service, and actively contribute to Agile workflows (e.g., demos, architecture sessions).
  • Create and maintain comprehensive technical documentation, including network diagrams, runbooks, and disaster recovery procedures to ensure system reliability and knowledge sharing.

Benefits

  • health, dental and vision insurance
  • 401(k)
  • flexible spending account
  • paid leave (including PTO and parental leave)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service