Engineering Manager, Edge SRE

Cloudflare•San Francisco, CA
•Onsite

About The Position

As part of the Cloudflare Infrastructure Engineering organization, our platform SREs are primarily responsible for production reliability. SREs are based in locations in Asia, Europe and the US enabling follow the sun coverage during daytime hours. SREs are supported by all engineering teams at Cloudflare who participate in on call schedules for their services. The SRE teams facilitate remediation and follow up of production issues and mature the tooling to enable all engineering teams to self-service on production. Incident follow up work across all engineering teams is prioritized above product innovation and the impact of production incidents influences the priority. SREs support two main environments: Edge SRE are focused on edge distribution where most client traffic is served. Core SRE are focused on the core services like control plane, data pipeline and other supporting supporting services. We are looking for an Engineering Manager to lead the Edge SRE team in London. You will lead and develop a team of SREs that are responsible for Cloudflare edge production and building the tools for all teams to understand and interact with it. You will play a lead role in driving our Platform initiatives for edge services and will be tasked with leading engineers who build tools and best practices for engineering teams to debug in production, measure availability and performance indicators, track and report on thresholds.

Requirements

  • 5+ years of software engineering, reliability, or operations experience in a customer-focused environment.
  • 2+ years experience managing a team of 5 or more engineers on projects in the areas of: distributed systems, tooling, Linux, Internetworking, infrastructure security or infrastructure management
  • Comfortable collaborating and co-ordinating on cross-team projects and workflows
  • Can provide a strong technical vision for systems and infrastructure teams
  • Experience building services and systems, have successfully taken projects from inception to production, and are comfortable diving in to provide leadership for major projects when needed
  • Capable of leading a discussion with upper management, and are able to tailor the level of technical detail to suit your audience

Nice To Haves

  • Hands-on experience with software or reliability engineering
  • Experience leading and hiring a team that builds and runs tools and platforms
  • Excel at planning and overseeing execution to meet commitments and deliver with predictability
  • Incident root cause analysis and follow-ups
  • Incident management
  • Comfortable managing teams/projections with deadlines and short release cycles
  • Experience using observability tools such as Jaeger, OpenTracing, ELK, Prometheus, Thanos, Grafana, Clickhouse
  • Experience running and maturing distributed systems
  • Familiarity working with Proxies, DNS, Databases, Internet and Security
  • Experience developing tools and APIs

Responsibilities

  • Lead a team of engineers who are working to keep the Cloudflare edge reliable and scalable
  • Mentor, grow, and empower your team by giving them the skills, confidence and motivation to make decisions
  • Help the individuals on your team to build and execute personal development plans that align with Cloudflare’s goals and objectives
  • Take an active role in prioritizing the roadmap for the SRE Org
  • Drive cross-team and cross-org alignment in engineering, infrastructure and product teams
  • Partner with other Engineering Managers across Cloudflare to achieve reliability outcomes for their services
  • Participate in deep technical design discussions within your team, and across partner teams, and ensure that we're building the right systems and keeping the quality high

Benefits

  • Project Galileo
  • Athenian Project
  • 1.1.1.1
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service