Staff Site Reliability Engineer - Kubernetes

OktaWashington, DC
$174,000 - $267,000Hybrid

About The Position

The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses on architecting and managing reliable, scalable, and secure Kubernetes-based platforms on AWS, ensuring high availability and performance while optimizing costs and automation. The ideal candidate will have hands-on experience with AWS infrastructure, Kubernetes platform creation, Helm charts, Karpenter scaling, and Istio service mesh. This role is part of the Workforce Identity Cloud team at Okta, which provides easy, secure access for organizations' workforces. The company is looking for individuals passionate about solving large-scale automation, testing, and tuning problems, who embody the principle of automating repetitive tasks and can rapidly self-educate on new concepts and tools.

Requirements

  • 4+ years of experience with Kubernetes/Helm
  • 4+ years of Experience with Terraform
  • 5+ years of Experience with AWS
  • Experience with multi-region cloud environments
  • Proven experience with AWS (EC2, RDS, S3, CloudFormation, IAM, etc.) and solid understanding of cloud-native architectures
  • Strong expertise in Kubernetes platform creation, management, and optimisation (e.g., setting up highly available clusters, networking, and storage)
  • Hands-on experience with Helm for Kubernetes application deployment and management
  • Practical experience with Karpenter for dynamic scaling of Kubernetes clusters and optimising resource usage
  • Expertise in managing and securing Istio for service mesh, including traffic management, security, and observability features
  • Proficiency in CI/CD pipelines and automation tools (e.g., Jenkins, GitLab, CircleCI, Terraform, Ansible, Spinnaker)
  • Strong scripting and automation skills in Python, Bash, or Go for infrastructure management and platform automation
  • Experience with monitoring, logging, and alerting tools such as Prometheus, Grafana, CloudWatch, and ELK Stack
  • Ability to access federal environments and/or have access to protected federal data
  • Must be able to submit documentation establishing U.S. Person status upon hire
  • Requires in-person onboarding and travel to San Francisco, CA HQ office or Chicago office during the first week of employment

Nice To Haves

  • Understanding of security best practices for cloud platforms and Kubernetes (e.g., role-based access control (RBAC), encryption, and compliance frameworks)
  • Familiarity with Docker and containerization principles
  • Bachelor’s degree in Computer Science, Engineering, or related field (or equivalent professional experience)
  • CKA (Certified Kubernetes Administrator), CKAD (Certified Kubernetes Application Developer), or AWS Certified DevOps Engineer certifications are highly desirable

Responsibilities

  • Design, implement, and maintain highly available, scalable, and fault-tolerant Kubernetes platforms, optimized for production workloads with high resilience and operational efficiency.
  • Build, manage, and optimize AWS cloud infrastructure, including EKS, ECS, S3, VPCs, RDS, IAM, and implement best practices for cost management, scaling, and security.
  • Utilize Helm to automate and streamline application and service deployment to Kubernetes clusters, including creating and managing Helm charts.
  • Implement and manage Karpenter for dynamic scaling of Kubernetes clusters based on workload demands.
  • Configure and manage Istio for service-to-service communication, security, and observability within Kubernetes clusters, enabling traffic management, service discovery, and policy enforcement.
  • Automate the deployment, scaling, and management of infrastructure and applications, working with CI/CD pipelines for seamless development to production transitions with minimal downtime.
  • Respond to incidents, troubleshoot, and resolve system issues related to performance, availability, and security in a timely and effective manner.
  • Design and implement secure cloud infrastructure with appropriate access controls, network security, and compliance frameworks.
  • Create and maintain detailed documentation for Kubernetes platform setup, operational procedures, and best practices, and promote knowledge sharing across teams.

Benefits

  • health, dental and vision insurance
  • 401(k)
  • flexible spending account
  • paid leave (including PTO and parental leave)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service