Site Reliability Engineering Lead

RELX•Washington, DC
•$118,300 - $219,800

About The Position

LexisNexis Risk Solutions is seeking a passionate and experienced Site Reliability Engineering Lead to join their Core SRE Team in Business Services. This role involves overseeing applications and infrastructure within a major business unit, building cloud environments, migrating on-prem applications to the cloud, and managing self-hosted and third-party solutions. The ideal candidate is a self-starter who can assess situations, develop collaborative solutions, and proactively improve performance, cost, and reliability. This is a professional management-level position requiring line management of a small to medium-sized team of engineers. Responsibilities include performance management, pay authority, recruitment, task prioritization, and providing support to team members. The lead will also support engineers' personal development, ensure adherence to the SRE framework, lead post-mortem reviews and RCA production, and address issues impacting multiple teams.

Requirements

  • Expert knowledge of Kubernetes, including cluster architecture, upgrades, autoscaling, security hardening, and troubleshooting at scale
  • Expert experience with Terraform, including modular IaC design, state management, multi-environment provisioning, and policy-as-code
  • Deep knowledge of Azure Cloud, including compute, networking, identity (AAD), storage, and cost optimization
  • Experience designing and scaling CI/CD pipelines using GitHub Actions, release strategies, and rollback automation
  • Experience with observability platforms including Prometheus, Grafana, OpenTelemetry, and SLO/SLA/error-budget management
  • Strong automation skills, focused on eliminating toil through self-healing systems and infrastructure automation
  • Advanced proficiency in Python, Bash, and/or PowerShell for tooling and automation
  • Deep understanding of networking concepts including TCP/IP, DNS, load balancing, VPNs, and cloud-native networking
  • Experience in SRE, DevOps, or Infrastructure roles, including experience leading engineering teams
  • Proven track record leading incident response and driving reliability improvements

Responsibilities

  • Manage, mentor, and grow a team of SREs; conduct 1:1s, performance reviews, and career development planning
  • Own hiring, onboarding, and team capacity/resourcing decisions
  • Set team goals, prioritize backlog, and drive planning
  • Foster a blameless post-incident culture and cross-team collaboration with Dev, Security, and Product
  • Lead reliability initiatives across infrastructure and services
  • Drive incident response activities and continuous service improvement
  • Champion automation and operational excellence across the platform
  • Support the development of scalable, secure, and resilient cloud-native environments

Benefits

  • Annual incentive bonus
  • Country specific benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service