Senior Software Engineer - SRE & AIOps

ServiceNowSanta Clara, CA
Remote

About The Position

ServiceNow is seeking a Senior Software Engineer - SRE & AIOps to contribute to infrastructure automation, operational resilience, and toil elimination across our hybrid cloud and data center operations. Embedded within the Site Reliability & Database Engineering organization, you will implement automation-first systems that reduce manual intervention, accelerate incident remediation, and enable our global engineering teams to operate reliably at scale. This role combines solid hands-on technical expertise in Kubernetes, cloud platforms, and DevOps practices with growing technical leadership capabilities. You will contribute to SRE tooling design, develop auto-remediation capabilities, and help establish patterns that maintain ServiceNow's cloud platform reliability while minimizing operational toil across follow-the-sun global teams.

Requirements

  • Kubernetes Proficiency: Solid hands-on experience operating production Kubernetes clusters, including deployment models, pod orchestration, resource management, network policies, and troubleshooting runtime issues.
  • Incident Remediation Experience: Demonstrated experience designing and implementing automated remediation systems, including alert automation, runbook development, and self-healing mechanisms.
  • Cloud Platform Knowledge: Strong hands-on experience with AWS (EKS, EC2, RDS) and/or Azure (AKS, VMs) or GCP (GKE), with understanding of core SRE-related services.
  • DevOps & IaC Skills: Solid experience with Infrastructure-as-Code tools (Terraform, CloudFormation) and GitOps practices.
  • SRE Tooling Familiarity: Working knowledge of observability platforms, incident management systems, and log aggregation tools.
  • Distributed Systems Understanding: Understanding of distributed system challenges, fault tolerance, and resilience patterns.
  • On-Call Operations: Experience participating in on-call rotations and understanding 24/7 operational models, runbook development, and escalation procedures.
  • Cloud & Hybrid Operations: Hands-on experience working with cloud infrastructure and understanding hybrid cloud concepts.
  • Systems Administration: Strong foundation in Linux system administration, performance troubleshooting, and scripting (Python, Go, or Bash).
  • Collaborative Mindset: Ability to work effectively with infrastructure and application teams, contribute to technical discussions, and help drive reliability improvements.
  • Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry.
  • 5+ years in software engineering or infrastructure operations, with 3+ years in SRE, DevOps, or cloud platform engineering roles with a Bachelor's degree; or 3 years and a Master's degree; or a PhD without experience; or equivalent work experience.
  • 2+ years of hands-on experience working with production Kubernetes clusters.
  • Proficiency in at least one Infrastructure-as-Code tool: Terraform, CloudFormation, or equivalent.
  • Demonstrable hands-on experience with at least one major cloud platform: AWS, Azure, or GCP.
  • Experience operating in on-call environments and participating in incident response.
  • Experience implementing or improving automated remediation and alert systems.
  • Strong foundation in Linux system administration, performance troubleshooting, and scripting (Python, Go, Bash).
  • Demonstrated commitment to reliability engineering and continuous improvement through hands-on contributions.
  • Bachelor's degree in computer science, Computer Engineering, or related field (or equivalent professional experience).

Nice To Haves

  • Kubernetes certification (CKA, CKAD, or equivalent).
  • Experience with service mesh technologies or advanced Kubernetes networking.
  • Background in cloud migration or infrastructure modernization projects.
  • Experience with cost optimization in cloud environments.
  • Track record of implementing automation solutions that significantly reduced operational toil.

Responsibilities

  • Deploy, operate, and troubleshoot production Kubernetes clusters across hybrid and multi-cloud environments, maintaining operational standards and supporting high-velocity application deployments.
  • Implement and maintain closed-loop auto-remediation systems that detect, classify, and resolve transient infrastructure failures, leveraging automation frameworks and machine learning insights to reduce MTTR and on-call burden.
  • Contribute to the design and evolution of SRE tooling stack, including monitoring platforms, incident management systems, log aggregation, and observability integrations that support global on-call operations.
  • Develop and maintain SLO frameworks, alerting policies, and automated runbooks that empower on-call engineers to resolve issues autonomously while managing alert fatigue.
  • Build and maintain Infrastructure-as-Code frameworks and GitOps pipelines that enable reproducible infrastructure deployments across hybrid and multi-cloud environments with security and compliance guardrails.
  • Support hybrid cloud and data center operations, including on-premises infrastructure, public cloud environments, and workload optimization across multi-region deployments.
  • Contribute to adoption of containerization, microservices, and DevOps patterns across engineering teams, establishing CI/CD best practices and network security controls.
  • Support on-call rotation operations and incident response processes across different time zones, helping develop runbooks and contributing to post-incident reviews that drive continuous improvement.
  • Share knowledge and mentor junior SRE engineers on reliability patterns, incident investigation techniques, and automation best practices.
  • Champion a culture of blameless incident analysis, data-driven decision-making, and continuous improvement through knowledge sharing and documentation.
  • Identify and systematically automate repetitive operational tasks, from infrastructure provisioning to incident response, improving team efficiency and capacity.

Benefits

  • health plans
  • flexible spending accounts
  • a 401(k) Plan with company match
  • ESPP
  • matching donations
  • a flexible time away plan
  • family leave programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service