We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary: About the Team Our Site Reliability Engineering team is the execution engine behind the reliability, availability, and performance of distributed store technology powering thousands of retail and pharmacy locations nationwide. We operate across pharmacy platforms, Point of Sale (POS) systems, handheld devices, store servers, dispensing systems, and edge computing infrastructure — spanning hybrid cloud and on-premises environments at massive fleet scale. Our engineering philosophy is grounded in five pillars: Detection, Prevention, Recovery, Learning Loops, and Developer Experience (DevX). Our operating principle is the reliability covenant: SRE's success is measured not by how many incidents we respond to, but by how much reliability capability we transfer to the engineering teams we serve. The ultimate measure of this function is how much of the work it currently does becomes unnecessary over time — because development teams have internalized reliability ownership, automation has replaced manual operations, and the organization has developed the reflexes to prevent rather than respond. If you are drawn to building a capability-transfer engine rather than an operations center, this is the right environment. We track operational toil as an engineering metric. Toil accumulation is treated as a reliability risk and a capacity cost. Engineers at every level are expected to eliminate it, document the reduction, and reinvest the recovered capacity into prevention and automation. About the Role As a Staff Software Engineer — SRE, you own programs and platforms that span all engineering domains. Where a Senior Engineer owns a domain's reliability posture, you own the infrastructure of reliability itself — the observability platform, the production readiness framework, the incident management operating model, the chaos engineering program, and the developer reliability platform. You author standards, not just follow them. You set the direction that SSE and SE engineers operate within, partner with engineering domain owners as a technical peer and organizational change agent, and represent SRE at the engineering director level. Your decisions have blast radius across the entire engineering organization and directly affect the operational experience of thousands of store locations. Scope: Cross-domain program ownership — you define the system, not just operate within it. The Environment You Are Joining This section is an honest description of what you are joining, not a caveat. Read it carefully. You are joining at the inception of an SRE transformation in a large-scale engineering organization. The programs described in this role — the observability platform, the chaos engineering program, the production readiness framework, the incident management operating model — are in early to mid stages of development. You will build them from first principles. The engineering organization you are partnering with has operated without a formalized SRE function for years. There are established ways of working, teams with strong ownership cultures, and leadership that is results-oriented and skeptical of new frameworks until they demonstrate measurable value. Adoption is earned, not assumed. You will need to demonstrate the value of every program you introduce before you can expect the organization to invest in it. The observability infrastructure you will architect does not yet exist at the maturity level described in the What You Will Do section. You will be building toward that state simultaneously with running operational programs. The first 90 days will involve as much organizational assessment and relationship-building as technical design. Candidates who thrive in this role are energized by the combination of deep technical architecture work, greenfield program building, and the organizational challenge of making a large, established engineering organization believe in something new. They have done this before — built programs from a blank sheet inside a complex organization — and they understand that the organizational work is as important as the technical work. Candidates who may find this role difficult are those who have operated in mature, well-defined SRE environments where the toolchains are established, the frameworks are proven, and adoption is already institutionalized. The ambiguity, the build mandate, and the organizational persuasion work are not temporary features of a ramp-up period — they are permanent features of the job at this stage of the function's development. What You Will Do Detection & Observability Architect the team's observability platform strategy end-to-end: hot-tier event streaming, warm-tier analytical query and cold-tier long-term storage — including schema design, data contracts, and retention policy Define the org-wide composite reliability signal framework: design how individual SLIs aggregate into domain health scores and fleet-wide reliability indicators that map to business outcomes — not just technical signals Drive SLI/SLO standardization across all engineering domains; own complete Critical User Journey (CUJ) coverage with documented gap analysis, business impact severity mapping, and a roadmap to close every gap Design and own the AI-enabled detection pipeline: implement time-series anomaly detection as a production system; integrate LLM-based alert summarization and incident triage assistance into the detection-to-response workflow; own the feedback loop that improves model accuracy over time from incident outcome data Design certificate and dependency lifecycle monitoring: automated discovery (via Vault, cert-manager, or PKI management API integrations), alerting thresholds with sufficient lead time, automated ticket creation, SRE approval gate, and closed-loop renewal validation Evaluate and adopt emerging observability standards including OpenTelemetry for unified telemetry collection and eBPF-based tooling for low-overhead infrastructure observability Design for the fleet: observability architecture must account for thousands of unattended edge nodes with intermittent connectivity — design for connectivity-tolerant telemetry collection, local buffering, and fleet-aggregated health signals that distinguish node-level failures from fleet-wide patterns Prevention & Reliability Engineering Own the Production Readiness Review (PRR) framework end-to-end: define the principles, author the checklist, build the YAML-based manifest or equivalent, and implement automated policy enforcement using OPA/Conftest, GitHub Actions gates, or custom CI/CD SRE review APIs — the goal is a gate that is frictionless for healthy services and informative for risky ones Design and run an organizational chaos engineering and GameDay program: define steady-state hypotheses, build failure injection scenarios using LitmusChaos or Chaos Toolkit, facilitate multi-domain GameDay exercises, and measure organizational readiness improvement across cohorts — not just exercise completion Lead fleet-scale deployment reliability: design progressive rollout strategies (canary cohorts, staged fleet expansion, geographic blast radius budgets) for edge node deployments where a misconfigured rollout can simultaneously affect thousands of unattended locations; own the deployment gate criteria and the rollback SLA framework Design configuration drift detection: implement systems that identify when deployed node configurations diverge from the expected state, alert before the drift causes incidents, and provide automated or semi-automated remediation paths Lead hardware and infrastructure risk assessments for major platform programs: evaluate compute sizing, memory constraints, and runtime isolation requirements at fleet scale before architectural commitments are made Define the organization's change management reliability gates: CHG review standards, rollback SLA requirements, deployment freeze windows aligned to peak operational periods, and blast radius assessment criteria for high-risk changes Partner with architecture review boards to embed reliability non-functional requirements (NFRs) into the design phase of major platform programs — reliability engineered in from the first whiteboard session, not reviewed at the launch gate Incident Response & Recovery Design and operationalize a tiered incident response model: define severity tiers, service and fleet criticality classifications, and proactive monitoring cohorts that enable SRE pre-engagement before customer impact at fleet scale Serve as a certified Technical Incident Commander (TIC); lead P0 and P1 incident bridges with cross-organizational stakeholder coordination spanning Engineering, Infrastructure, Operations, and executive leadership Own the on-call operating model: design the shift structure (follow-the-sun coverage, domain-assigned ownership tiers), conduct coverage gap analysis, define escalation SLAs, and run quarterly on-call health reviews that surface toil concentration, coverage gaps, and alert quality trends Design the alert management and suppression governance framework: suppression policies for planned maintenance and deployments, alert routing rules, and criteria for alert promotion/demotion across severity levels — operating within tools such as Grafana IRM, PagerDuty, or OpsGenie Own the organization's MTTR reduction program as a measurable engineering initiative; establish the MTTR-within-SLA baseline and drive systematic improvement through runbook quality, automation, and escalation path optimization — not heroics Learning Loops & Continuous Improvement Learning Loops at this level is data pipeline engineering for organizational memory — not postmortem program management. Own the postmortem program end-to-end: design the causal taxonomy (origin layer classification + failure pattern codes as a queryable data structure), define the quality scoring model, build the facilitation framework, and run the review cadence across all incident types and severity levels Design the three-phase incident learning model: Phase 1 — structured operational record (AI-assisted scribe using LLM tooling for real-time transcription and timeline reconstruction); Phase 2 — lightweight root cause analysis authored by the TIC within 24 hours; Phase 3 — full blameless postmortem facilitated jointly with Problem Management Engineer the feedback loop as a data pipeline: incident data → causal taxonomy classification → PRR gate update triggers → development team practice changes → reduced recurrence. Own the metrics that prove the pipeline is working: incident class recurrence rate, time-from-incident-to-PRR-update, and reduction in repeat incident classes by causal origin Measure learning velocity as an organizational metric: the average time from first occurrence of an incident class to systemic resolution — not just runbook coverage but actual recurrence elimination Partner with the Problem Management organization to align postmortem quality standards, systemic issue escalation criteria, and cross-organizational incident learning governance Own the P1/P2 trend analysis: identify systemic incident drivers across all domains, quantify their business impact, and build the investment case for engineering programs that break the repeat-incident cycle Developer Experience & Automation (DevX) Own the SRE Developer Platform: design and build the internal developer tooling library including SRE onboarding kits for new service development (greenfield patterns), observability integration templates (brownfield patterns), developer-facing reliability APIs, and self-service SLO health pages Define the organization's DORA metrics program and establish org-wide baselines: instrument Deployment Frequency, Lead Time for Change, MTTR, and Change Failure Rate — run monthly reporting to engineering leadership and connect the data to reliability investment decisions Drive self-service reliability adoption: developer-owned SLO dashboards, automated PRR manifests triggered by CI/CD events, and service health pages that surface CUJ health to the owning team — the measure of success is not adoption rate, it is the number of reliability decisions engineering teams can make without filing an SRE ticket Evangelize the reliability covenant across engineering domain owners: frame SRE as a capability transfer function, not a gatekeeper or an operations team. Your success is the engineering team's independence, not your indispensability Quantify and report toil economics at the organizational level: total manual operational touchpoints per month across all domains, automation coverage rate, and the dollar value of capacity recovered through toil elimination — use this data to make the case for continued investment in automation and developer platform tooling Organizational Influence & Change Management Build organizational buy-in before building systems: before a program is widely adopted, it must be widely believed in. Your ability to make engineering directors, domain owners, and platform architects understand why a program matters — and what it will take from them — is as important as your ability to design the program itself Develop and execute an adoption strategy for each major program you own: identify the minimum viable version that delivers clear value to a skeptical team, land it, measure the outcome, and use the evidence to expand adoption across the organization Diagnose and manage organizational resistance: when a program is not being adopted, identify the real reason — unclear value proposition, wrong timing, too much friction, competing priorities — and adapt the approach rather than escalating to mandate Establish and maintain relationships with engineering domain owners as a peer, not a gating function; understand their delivery pressures, quality concerns, and what they need from SRE to say yes to reliability investments Represent SRE at the engineering leadership level: communicate program health, reliability trends, investment priorities, program ROI in terms of business outcomes — not SRE metrics
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior