Staff Site Reliability Engineer, EPG (FedRAMP)

OktaWashington, DC
$174,000 - $267,000Hybrid

About The Position

Okta is seeking an experienced Staff Site Reliability Engineer to join their Emerging Products Group (EPG). The team's mission is to build highly reliable, scalable, and secure cloud services. They emphasize an automation-first mindset and invest in platform engineering, observability, and operational excellence. This role is for a seasoned engineer who enjoys solving complex, cross-team technical challenges and will serve as a key technical lead. The ideal candidate embodies the philosophy of automating repetitive tasks and has a strong passion for continuous improvement, operational excellence, and software engineering.

Requirements

  • Extensive experience architecting and leading the evolution of large-scale production services in AWS and/or GCP.
  • Deep expertise in defining Kubernetes patterns and Linux-based system standards for enterprise-grade production environments.
  • Experience designing multi-region, highly available cloud architectures from the ground up.
  • Experience troubleshooting Kubernetes networking, storage, scheduling, scaling, and workload lifecycle issues.
  • Proven track record of evaluating "build vs. buy" decisions and setting long-term technical standards for an organization.
  • Extensive experience with Infrastructure as Code technologies such as Terraform and Helm.
  • Strong software engineering skills in Golang and/or Python.
  • Experience building automation and internal engineering platforms.
  • Experience operating and troubleshooting distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, Cassandra, or similar technologies.
  • Strong understanding of cloud networking fundamentals including DNS, load balancing, ingress, TLS, service networking, and traffic management.
  • Strategic experience designing comprehensive observability frameworks and telemetry-driven operational strategies.
  • Experience with or strong interest in AI-assisted engineering and operational automation.
  • Strong expertise operating customer-facing production systems subject to SLA.
  • Experience leading incident response and driving operational improvements.
  • Deep understanding of reliability engineering concepts including SLIs, SLOs, error budgets, and capacity planning.
  • Strong understanding of CI/CD pipelines, deployment strategies, and automation-first operational practices.
  • Proven ability to balance reliability, scalability, security, and engineering velocity.
  • Understanding of cloud security fundamentals, IAM, secrets management, and secure infrastructure design.
  • Demonstrated success leading complex engineering initiatives across multiple teams.
  • Strong collaboration and communication skills.
  • Experience working effectively within globally distributed engineering organizations spanning multiple timezones and cultures.
  • Act as a technical force-multiplier, defining engineering standards, raising the technical bar across the EPG organization, and sponsoring the growth of junior engineers.
  • Ability to influence technical direction through expertise, partnership, and execution.
  • This role supports US FedRAMP projects and requires the employee to be a US Person (US Citizen or Green Card Holder) to meet FedRAMP compliance and security clearance standards.
  • The Okta employee must be on US Soil, which means the 50 states, the District of Columbia, or outlying areas of the United States, as defined in Federal Acquisition Regulation (FAR) 2.101.

Nice To Haves

  • Experience implementing operational controls, compliance standards (e.g., FedRAMP, SOC2, HIPAA), and best practices in highly regulated or security-sensitive government cloud environments is highly preferred.
  • Experience operating SaaS platforms serving large-scale customer workloads.
  • Experience working within Kubernetes-based microservices environments.
  • Experience supporting globally distributed production environments.
  • Experience with GitOps and ArgoCD.
  • Experience implementing AI-assisted operational tooling or automation workflows.

Responsibilities

  • Design, build, and operate large-scale cloud infrastructure and production services.
  • Participate in a global on-call rotation supporting highly available customer-facing systems.
  • Lead incident response efforts and drive post-incident reviews focused on systemic improvements.
  • Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Partner with engineering teams to improve service availability, scalability, performance, and resilience.
  • Ensure all infrastructure and operational practices adhere to strict FedRAMP compliance and security mandates, maintaining continuous audit readiness.
  • Continuously improve observability through metrics, logging, tracing, dashboards, and alerting.
  • Develop software, automation, and infrastructure using Go, Python, Terraform, and related technologies.
  • Eliminate operational toil through automation, tooling, and platform engineering.
  • Improve deployment safety and operational workflows through CI/CD and GitOps practices.
  • Collaborate on modernizing existing workloads and aligning them with evolving platform capabilities.
  • Build self-service platforms, operational guardrails, and automation that improve developer velocity while maintaining reliability and security.
  • Lead complex reliability initiatives spanning multiple engineering teams.
  • Guide engineers in adopting operational best practices and reliability engineering principles.
  • Mentor engineers through technical collaboration, design reviews, incident analysis, and knowledge sharing.
  • Influence architecture and operational decisions through data-driven recommendations and engineering expertise.
  • Drive projects from conception through production rollout and long-term operational ownership.
  • Explore and apply AI-assisted engineering techniques to improve operational efficiency, incident response, troubleshooting, and automation.
  • Identify opportunities to leverage emerging technologies to reduce toil and improve engineering productivity.

Benefits

  • health, dental and vision insurance
  • 401(k)
  • flexible spending account
  • paid leave (including PTO and parental leave)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service