Staff Reliability Engineer

ServiceNowSanta Clara, CA
$166,500 - $291,400Hybrid

About The Position

Join us to build the next generation of cloud-native reliability, release, and test platforms that enable engineering excellence, developer productivity, and high-confidence ServiceNow releases through automation, observability, and AI-driven operations.

Requirements

  • Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry.
  • 8+ years of experience in Site Reliability Engineering (SRE), DevOps, Platform Engineering, Software Engineering, or Infrastructure Engineering with a Bachelor's degree; or 6 years and a Master's degree; or a PhD with 3 years experience; or equivalent experience.
  • Hands-on experience with Kubernetes across cluster operations, networking, storage, security, autoscaling, and multi-cluster environments.
  • Experience building and operating cloud-native platforms supporting scalable, highly available services.
  • Experience integrating Kubernetes with CI/CD, GitOps, automated test pipelines, deployment validation, and cloud-native deployment workflows.
  • Experience designing and implementing automation to improve developer productivity, release quality, and operational efficiency.
  • Experience with progressive delivery practices, including canary deployments, feature flags, automated rollback, and deployment verification.
  • Experience with chaos engineering, resilience testing, disaster recovery, and reliability validation.
  • Strong software engineering skills with hands-on experience designing, developing, testing, and debugging applications using Python, Go, Java, or Ruby.
  • Experience leveraging AI-assisted engineering for intelligent testing, release risk analysis, incident diagnostics, or operational automation is a plus.
  • Strong understanding of observability, monitoring, SLI/SLOs, incident management, and production operations for distributed systems.
  • Demonstrated ability to solve complex technical problems, drive projects independently, and collaborate effectively across engineering teams.
  • Thrives in fast-paced, ambiguous environments with a strong ownership mindset, bias for action, and a passion for continuous learning and automation.
  • Low ego, intellectually curious, and an effective collaborator who enjoys partnering with globally distributed teams to deliver reliable engineering solutions.

Nice To Haves

  • Experience with observability and monitoring platforms for applications, services, and distributed systems at scale.
  • Experience with DevOps automation, CI/CD pipelines, GitOps, and Agile development practices using tools such as GitLab CI/CD, Argo CD, or Flux.
  • Experience building and maintaining enterprise-scale test automation frameworks using technologies such as Playwright, Selenium, Cypress, REST Assured, PyTest, JUnit/TestNG, or equivalent.
  • Experience with test orchestration, intelligent regression testing, test impact analysis, flaky test detection, parallel execution, and test data management.
  • Experience with service virtualization, contract testing, synthetic testing, and building developer self-service engineering platforms.
  • Experience with Infrastructure as Code and configuration management tools such as Ansible, Terraform, or equivalent.
  • Experience with the Kubernetes ecosystem, including Helm, Argo Workflows, Kustomize, Istio/Linkerd, Gateway API/Ingress, Prometheus, OpenTelemetry, and container runtime technologies.
  • Experience operating Kubernetes platforms across public cloud providers, including AWS (EKS), Azure (AKS), and Google Cloud (GKE).
  • Experience implementing progressive delivery practices, including canary deployments, feature flags, deployment verification, and automated rollback.
  • Familiarity with AI-assisted engineering, intelligent testing, operational automation, or cloud-native engineering platforms.

Responsibilities

  • Design, build, and operate cloud-native engineering platforms for software validation, release validation, and production readiness
  • Design and maintain production-like release and test ServiceNow environments that improve release confidence and deployment readiness.
  • Build and integrate automated test pipelines, observability, reliability signals, deployment intelligence, and quality gates into CI/CD workflows.
  • Develop automation solutions that improve engineering productivity, streamline operations, and reduce manual toil through shift-left engineering practices.
  • Build reusable frameworks, self-service engineering environments, test data management, mock services, and developer productivity tooling.
  • Design and enhance Kubernetes-based platforms supporting scalable test infrastructure, release automation, cloud-native workloads, and developer self-service.
  • Implement automated validation for failure detection, deployment verification, policy enforcement, security checks, resilience testing, and operational health assessments.
  • Resolve complex platforms, infrastructure, and networking challenges through software engineering, systems design, and automation.
  • Partner closely with engineering teams to improve platform reliability, release quality, cloud-native adoption, and engineering best practices.
  • Participate in architecture reviews, technical design discussions, and implementation of scalable, automation-first engineering solutions.
  • Influence technical decisions through strong engineering execution, collaboration, and delivery of high-quality platform capabilities.
  • Mentor engineers through technical guidance, code reviews, knowledge sharing, and engineering best practices.
  • Foster a culture of reliability, automation, operational excellence, continuous improvement, and customer-focused engineering.

Benefits

  • health plans
  • flexible spending accounts
  • a 401(k) Plan with company match
  • ESPP
  • matching donations
  • a flexible time away plan
  • family leave programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service