Lead Site Reliability Engineer

Lumen Technologies,
$105,786 - $155,152Remote

About The Position

Lumen is the trusted network for the AI‑powered world, connecting people, data, and applications through our expansive fiber network and connected ecosystem. We enable secure, high‑performance connectivity across cloud, edge, and AI workloads for enterprises, governments, and communities. At Lumen, you’ll work on infrastructure customers rely on today and build for what’s next, where performance, security, and resilience matter. This is a high accountability environment where bold ideas drive real innovation for our customers, partners, and industry. The work is challenging, expectations are clear, and trust is built into how we operate. If you’re ready to take ownership, deliver meaningful impact, and help shape the future of AI‑ready connectivity, join us today. Lumen is seeking a Lead Site Reliability Engineer (SRE) who will be a catalyst for transformational change and operational excellence. The ideal candidate not only builds and operates systems but proactively identifies gaps, challenges legacy approaches, and delivers high-impact changes that improve customer experience. You will own platform engineering standards across Azure Public and Azure Government clouds, lead the design of CI/CD and observability systems, and partner closely with development teams to deliver scalable, secure, and compliant high-availability services.

Requirements

  • 8+ years in DevOps/Platform/SRE roles; 2+ years leading teams or cross-functional initiatives.
  • Deep expertise in Azure (Public/Gov) cloud services and resource management.
  • Hands-on with AKS, Helm 3, Flux CD (GitOps), Docker.
  • Experience with Prometheus & Grafana (metrics, alerting, dashboards).
  • Proficiency in Nginx and Kong (ingress/API gateway).
  • CI/CD automation using Jenkins and GitHub Actions.
  • Strong development skills for automation & debugging (Python, Go, or TypeScript; Bash/PowerShell).
  • Proven experience with IaC using Terraform (or similar tools like Bicep), policy-as-code, and secure SDLC integrations.
  • Ability to obtain GSA Tier 2 Suitability Clearance.

Nice To Haves

  • Experience in Azure Government with FedRAMP controls and ATO support.
  • Certifications: AZ-104, AZ-305, AZ-400, CKA/CKAD, Terraform Associate.

Responsibilities

  • Design, build, and operate AKS clusters with enterprise guardrails (RBAC, pod security policies, node pools, autoscaling, upgrade strategies).
  • Implement GitOps with Flux CD; define Helm chart standards and manage lifecycle via Helm 3.
  • Own container platform best practices: Docker image standards, multi-stage builds, SBOM, vulnerability scanning, and base image governance.
  • Configure Nginx (ingress controller) and Kong (API gateway) including routing, rate limiting, mutual TLS, JWT/OAuth2 auth, and WAF integration.
  • Stand up Prometheus for metrics scraping, alert rules, and exporters; integrate Grafana for dashboards and SLO/SLA visualization.
  • Lead Jenkins pipeline automation and GitHub Actions workflows for build/test/deploy, with reusable templates and environment promotion strategies.
  • Build infrastructure-as-code and policy-as-code pipelines using Terraform (or similar tools like Bicep), with automated validations, PR gates, and compliance checks.
  • Architect and enforce secure integration with Azure Key Vault for secret, key, and certificate management; automate cert-manager/issuer (ACME/PFX).
  • Design and implement Azure Networking: VNets, NSGs, route tables, Private Link, Firewall rules, egress control, and traffic segregation in hub-spoke models.
  • Operate and integrate Azure Blob Storage, Azure Redis Cache, PostgreSQL, Azure Cosmos DB (Mongo API), and Azure Service Bus.

Benefits

  • Health, Life, Voluntary Lifestyle benefits and other perks that enhance your physical, mental, emotional and financial wellbeing.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service