Director, Site Reliability Engineering

SalesforceDallas, NY
Hybrid

About The Position

Salesforce is seeking a Director of Site Reliability Engineering to lead the advancement of our reliability, observability, and operational engineering capabilities. This role involves transforming the SRE function to shift the engineering organization from reactive incident response to a proactive, automated, and data-driven reliability culture. The Director will collaborate closely with Application Engineering, Platform, Architecture, Security, Infrastructure, and Product teams to ensure services are resilient, observable, scalable, and production-ready prior to launch. This is an impactful people leadership role requiring sharp technical judgment, managing a team of approximately 6 engineers, and driving cross-functional alignment within a complex organization. The individual will be responsible for defining the strategy, tooling, automation, and culture necessary to mentor their team and scale system reliability across the enterprise.

Requirements

  • Bachelor’s degree in Computer Science, Computer Engineering, Software Engineering, or a related technical field; Master’s degree or MBA preferred.
  • 10+ years of progressive engineering experience, including 5+ years in engineering leadership managing SRE, Platform, or Systems Engineering teams
  • Proven experience building or transforming a reliability or operational engineering organization.
  • Strong understanding of distributed systems, cloud architecture, application architecture, networking, infrastructure, and software delivery.
  • Experience establishing observability, incident-management, service-level objective, and production-readiness practices.
  • Demonstrated ability to improve reliability through engineering and automation rather than process alone.
  • Proven experience leading teams responsible for highly available, customer-facing, or business-critical systems.
  • Strong understanding of modern telemetry, including metrics, logs, distributed tracing, synthetic monitoring, and real-user monitoring.
  • Demonstrated success driving alignment and building consensus across cross-functional engineering teams and executive stakeholders
  • Ability to balance immediate operational needs with long-term engineering transformation.
  • Strong written, verbal, and executive communication skills.

Nice To Haves

  • Strong experience operating large-scale systems in AWS or another major cloud environment.
  • Proven track record with observability platforms such as New Relic, Splunk, Datadog, Sentry, Honeycomb, Grafana, Prometheus, or OpenTelemetry.
  • Demonstrated experience implementing OpenTelemetry or common instrumentation standards.
  • Verified proficiency building internal developer platforms, paved roads, or self-service reliability capabilities.
  • Experience applying AI, machine learning, or agent-based automation to operational workflows.
  • Seasoned capability with chaos engineering, resilience testing, disaster recovery, capacity planning, and performance engineering.
  • Software engineering experience and the ability to engage deeply in architecture and design discussions.
  • Solid background in supporting high-profile launches, events, or systems with significant customer and business impact.

Responsibilities

  • Define and execute the long-term strategy and roadmap for Site Reliability Engineering.
  • Establish a clear operating model for SRE, including team scope, engagement models, ownership boundaries, and success measures.
  • Build and develop a high-performing team of site reliability and operations engineers.
  • Modernize the SRE function through automation, AI-assisted operations, self-service capabilities, and engineering-first practices.
  • Translate business priorities and customer impact into clear reliability investments and engineering outcomes.
  • Advise senior technology leaders on operational risk, resilience, capacity, and reliability tradeoffs.
  • Establish service-level indicators, service-level objectives, error budgets, and reliability standards for critical services.
  • Partner with engineering teams to design reliability, scalability, recoverability, and graceful degradation into systems.
  • Define what it means for a service to be operationally and observably ready for production.
  • Develop readiness reviews and certification practices for high-impact services and launches.
  • Drive improvements in system availability, performance, resiliency, and recovery.
  • Ensure reliability requirements are incorporated throughout the software development lifecycle rather than addressed only after deployment.
  • Define an enterprise observability strategy spanning metrics, logs, traces, events, synthetics, real-user monitoring, and business telemetry.
  • Establish common instrumentation, telemetry, dashboards, alerting, and service-health standards.
  • Reduce fragmented or duplicative observability implementations by promoting shared patterns and reusable capabilities.
  • Improve end-to-end visibility across distributed systems, customer journeys, services, and infrastructure.
  • Partner with engineering teams to ensure telemetry is actionable, contextual, and tied to customer and business outcomes.
  • Establish governance and measurement to assess adoption and effectiveness of observability standards.
  • Improve incident detection, response, mitigation, communication, and learning.
  • Lead the transition from manual and reactive operations toward automated detection, diagnosis, remediation, and incident creation.
  • Reduce mean time to detect, acknowledge, mitigate, and recover.
  • Improve on-call practices, escalation paths, runbooks, and operational ownership.
  • Establish blameless post-incident review practices that produce measurable engineering improvements.
  • Identify recurring sources of operational toil and create plans to eliminate or automate them.
  • Partner with engineering leaders to ensure actions from incidents are prioritized and completed.
  • Develop a roadmap for intelligent operations, including anomaly detection, event correlation, automated triage, assisted root-cause analysis, and remediation.
  • Evaluate opportunities to use agents and AI-assisted workflows across observability, incident response, capacity planning, and operational support.
  • Build automation that reduces cognitive load and improves the speed and consistency of operational decisions.
  • Ensure automation is safe, measurable, auditable, and designed with appropriate human oversight.
  • Promote platform and self-service approaches that allow product teams to adopt reliability practices with minimal friction.
  • Partner with engineering, DevOps, and business stakeholders.
  • Influence teams that do not directly report into SRE and build shared accountability for production outcomes.
  • Create clear service ownership models and operational expectations across teams.
  • Support major launches and critical business events through readiness planning, risk assessment, testing, and operational coordination.
  • Communicate reliability posture, risks, trends, and investments to executive and technical audiences.

Benefits

  • time off programs
  • medical
  • dental
  • vision
  • mental health support
  • paid parental leave
  • life and disability insurance
  • 401(k)
  • employee stock purchasing program
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service