Senior Staff Software Engineer – SRE, Release & Test Platforms

ServiceNowSanta Clara, CA
$190,900 - $334,100Hybrid

About The Position

Join us to build the next generation of cloud-native reliability, release, and test platforms that enable engineering excellence, developer productivity, and high-confidence ServiceNow releases through automation, observability, and AI-driven operations.

Requirements

  • Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry.
  • 12+ years of experience in software, systems, platform, or reliability engineering with a Bachelor's degree; or 8 years and a Master's degree; or a PhD with 5 years experience; or equivalent experience.
  • Deep Kubernetes expertise across architecture, operations, networking, storage, security, autoscaling, and multi-cluster environments.
  • Experience building and operating large-scale Kubernetes platforms for cloud-native, mission-critical services.
  • Experience integrating Kubernetes with CI/CD, GitOps, automated testing, and deployment validation.
  • Experience designing cloud-native platforms for ephemeral environments, release qualifications, and automated validation.
  • Proven ability to lead engineering excellence across developer productivity, platform engineering, release confidence, and modernization.
  • Experience with progressive delivery, including canary releases, feature flags, automated rollback, and deployment verification.
  • Experience with chaos engineering, resilience validation, disaster recovery, and reliability assessments.
  • Expertise designing, authoring, testing, and debugging code in a team setting using languages such as Python, Go, Java, or Ruby.
  • Experience using AI-assisted engineering for intelligent testing, release risk analysis, incident diagnostics, and operational automation.
  • Strong coding, observability, SLO, and cross-team collaboration skills to improve reliability, performance, and engineering standards.

Nice To Haves

  • Expertise in observability and monitoring applications, services, and networks at scale.
  • Experience with DevOps automation, CI/CD pipelines, and agile methodologies, including GitLab CI/CD or similar tools.
  • Experience building enterprise-scale test automation frameworks such as Playwright, Selenium, Cypress, REST Assured, PyTest, JUnit/TestNG, or equivalent technologies.
  • Experience with test orchestration, test impact analysis, flaky test detection, parallel execution, and intelligent regression testing.
  • Experience with service virtualization, contract testing, synthetic testing, and test data management.
  • Experience building engineering platforms that support developer self-service and release engineering.
  • Experience with infrastructure configuration management tools such as Ansible.
  • Expertise with Kubernetes ecosystem technologies such as Helm, Argo CD, Argo Workflows, Kustomize, Istio/Linkerd, Gateway API/Ingress, Prometheus, OpenTelemetry, and container runtimes.
  • Experience implementing GitOps using Argo CD, Flux, or similar technologies.
  • Experience operating Kubernetes across AWS (EKS), Azure (AKS), and Google Cloud (GKE).

Responsibilities

  • Build and operate cloud-native engineering platforms for software validation, release qualification, and operational readiness.
  • Design production-like release and test environments that improve release confidence and deployment readiness.
  • Develop automated quality gates to assess release health, operational risk, and production readiness.
  • Integrate automated testing, observability, reliability signals, and deployment intelligence into CI/CD pipelines.
  • Build reusable test frameworks, self-service environments, test data, mock services, and developer productivity tooling.
  • Advance shift-left engineering through automated validation, continuous verification, and quality gates.
  • Automate failure detection, policy validation, deployment verification, security checks, and reliability assessments.
  • Lead Kubernetes-based platform evolution for scalable test infrastructure, release automation, and developer self-service.
  • Resolve recurring infrastructure issues through sustainable software, systems, and networking solutions.
  • Partner with engineering teams on design reviews, architecture standards, and automation-first reliability practices.

Benefits

  • health plans
  • flexible spending accounts
  • a 401(k) Plan with company match
  • ESPP
  • matching donations
  • a flexible time away plan
  • family leave programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service