Senior Site Reliability Engineer

AkamaiCambridge, MA
Hybrid

About The Position

The AI Hardware SRE team is responsible for overseeing, scaling, and optimizing our next-generation dedicated AI hardware infrastructure. You will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings. This position focuses on enhancing system reliability, scalability, and performance across high-density hardware and software infrastructure in regional data centers. Responsibilities include defining KPIs, proactive monitoring, automation, and urgent issue resolution. Collaboration with teams ensures best practices, reduced downtime, and optimized systems. The role supports seamless operations, business-critical applications, and improved user experiences through efficient, data-driven solutions.

Requirements

  • 5+ years of relevant experience and a Bachelor's degree in Computer Science or related field
  • Exceptional proficiency in tooling and coding using languages like Python to build scalable operational tools, API integrations, and automation frameworks.
  • Hands-on experience with modern observability stacks and timeseries engines, like Prometheus, Grafana, OpenTelemetry, and Loki.
  • A working understanding of advanced networking topologies, high-bandwidth routing/switching infrastructure, BGP, and dual-stack IPv4/IPv6 networks.
  • Expertise as a primary designer for new service rollouts, establishing operational readiness criteria, telemetry baselines, and alerting thresholds.
  • Extensive experience building technical runbooks, leading complex incident response bridges, and driving comprehensive, blameless post-mortems.
  • A proven ability to fully own ambiguous technical challenges, coordinate cross-functional teams, and drive toward production-grade solutions effectively.

Responsibilities

  • Developing and scaling robust programmatic tooling and infrastructure-as-code utilities in Python to eliminate operational toil and automate fleet-wide provisioning.
  • Integrating automated workflows across diverse corporate ticketing systems to enhance resolution times for hardware and network break-fix incidents.
  • Leveraging advanced AI tools and LLM-based development approaches to enhance technical execution, script creation, and comprehensive system evaluation.
  • Working on cutting-edge private cloud and compute technologies to improve the availability, latency, and overall systemic health of high-density hardware environments.
  • Designing and implementing telemetry pipelines, custom Prometheus/Grafana monitoring dashboards, and AI-based anomaly detection tailored for bare-metal and virtualized environments.
  • Participating in 24x7x365 on-call rotations, spearheading real-time incident management, and managing high-severity service disruption protocols via automated PagerDuty and Slack workflows.
  • Partnering directly with third-party infrastructure vendors and coordinating on-site field technicians to facilitate uptime activities

Benefits

  • We support your health, well-being, finances, and life beyond work.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service