Site Reliability Engineering Lead

RELX•Alpharetta, GA
•$118,300 - $219,800

About The Position

LexisNexis Risk Solutions is seeking a Site Reliability Engineering Lead to provide technical leadership and strategic direction for Site Reliability Engineering initiatives across multiple product portfolios. This role involves leading a team of SREs responsible for ensuring the reliability, scalability, security, performance, and operational excellence of mission-critical applications and platforms. The SRE Lead will collaborate closely with engineering, architecture, security, operations, and business stakeholders to drive cloud modernization, operational maturity, observability excellence, automation, and continuous improvement. The position combines hands-on technical expertise with people leadership, mentoring, strategic planning, and cross-functional collaboration.

Requirements

  • Strong leadership experience managing technical engineering teams.
  • Deep expertise in Site Reliability Engineering, DevOps, Cloud Engineering, or Platform Engineering disciplines.
  • Extensive experience with Azure and/or AWS cloud platforms.
  • Strong understanding of Kubernetes, AKS, EKS, containerization, Docker, and cloud-native architectures.
  • Expertise with Infrastructure as Code tools such as Terraform and Ansible.
  • Strong background in observability platforms such as Grafana, Prometheus, OpenTelemetry, Splunk, Dynatrace, Datadog, or similar technologies.
  • Experience managing large-scale production environments with stringent availability requirements.
  • Strong understanding of security, compliance, networking, and cloud governance principles.
  • Experience designing highly available, fault-tolerant, and resilient systems.
  • Strong proficiency in at least one scripting or programming language such as Python, Go, PowerShell, Bash, or C#.
  • Experience with CI/CD pipelines and software delivery automation.
  • Exceptional troubleshooting and problem-solving capabilities.
  • Excellent communication and stakeholder management skills.
  • Strong documentation and presentation skills.
  • Ability to influence technical direction across multiple engineering organizations.
  • 8+ years of experience in Cloud Engineering, DevOps, Platform Engineering, Infrastructure Engineering, or Site Reliability Engineering.
  • 2+ years of leadership or people management experience leading engineering teams.
  • Bachelor's degree in Computer Science, Engineering, Information Systems, or equivalent practical experience.
  • Proven track record leading reliability and operational excellence initiatives in large-scale enterprise environments.

Nice To Haves

  • Experience building and managing enterprise-scale observability platforms.
  • Knowledge of FinOps, cloud cost optimization, and operational efficiency practices.
  • Experience with secret management platforms such as HashiCorp Vault, Akeyless or cloud-native alternatives.
  • Familiarity with Chaos Engineering and resilience testing.
  • Experience supporting regulated environments and compliance frameworks.
  • Experience leading cloud transformation and modernization programs.
  • Agile and Lean delivery experience.
  • Experience supporting global, distributed engineering teams.
  • Azure, AWS, Kubernetes, Terraform, or related certifications preferred.

Responsibilities

  • Lead and mentor a team of Site Reliability Engineers, fostering a culture of ownership, operational excellence, collaboration, and continuous learning.
  • Define and drive SRE strategy, standards, best practices, and operational frameworks across engineering organizations.
  • Partner with product and platform teams to improve application reliability, scalability, security, performance, and resilience.
  • Establish and maintain service level objectives (SLOs), service level indicators (SLIs), and error budgets.
  • Lead major incident management, root cause analysis, problem management, and post-incident review processes.
  • Drive cloud modernization initiatives and support application migrations to Azure, AWS, and containerized environments.
  • Champion automation and Infrastructure as Code (IaC) practices using tools such as Terraform, GitHub, GitLab, Jenkins, and Ansible.
  • Develop and implement observability strategies utilizing metrics, logs, traces, alerting, and dashboards.
  • Collaborate with security and compliance teams to ensure platform adherence to enterprise security and regulatory requirements.
  • Lead architecture reviews and provide guidance on cloud-native and highly resilient application designs.
  • Drive capacity planning, performance optimization, cost management, and operational efficiency initiatives.
  • Establish engineering guardrails, governance controls, and deployment standards for production environments.
  • Support organizational transformation toward DevOps and SRE practices.
  • Manage operational risk and ensure business continuity and disaster recovery preparedness.
  • Collaborate with stakeholders to prioritize reliability improvements and platform investments.
  • Build and maintain strong relationships with product owners, engineering leaders, vendors, and business partners.
  • Lead, coach, mentor, and develop a high-performing team of Site Reliability Engineers.
  • Conduct resource planning and support hiring, onboarding, and career development activities.
  • Establish team objectives aligned with business and technology strategies.
  • Promote accountability, innovation, and operational excellence within the team.
  • Act as a trusted advisor and subject matter expert for reliability engineering across the organization.
  • Drive cross-team collaboration and alignment on strategic initiatives.

Benefits

  • annual incentive bonus
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service