Senior II Site Reliability Engineer

AkamaiCambridge, MA
$146,400 - $263,600Hybrid

About The Position

This role is for a Senior II Site Reliability Engineer on the Akamai Inference Cloud Team, which is part of Akamai's Cloud Technology Group. The team designs, implements, deploys, and operates AI platforms for customer inference models and developer AI applications. In this position, you will lead reliability workstreams for Akamai's serverless inference platform, design SRE tooling and automation, and drive technical decisions. There are opportunities to mentor other SREs, influence architecture decisions with product engineering teams, and shape SRE practices for AI inference workloads and GPU infrastructure at scale.

Requirements

  • 8+ years of experience in SRE, infrastructure engineering, or platform engineering, working with large-scale distributed systems
  • Proven track record of defining SLO/SLI frameworks, building observability platforms, and running incident management processes at scale
  • Extensive Kubernetes and containerization experience at scale, including autoscaling, resource scheduling, and container orchestration for compute-intensive workloads
  • Experience building automation and tooling in Python or Go, with familiarity in CI/CD pipelines, deployment safety, and infrastructure-as-code
  • Ability to lead technical initiatives across teams, mentor other engineers, and drive complex reliability problems to resolution independently
  • Experience with or exposure to AI/ML infrastructure, model serving, or GPU workloads

Responsibilities

  • Taking ownership of observability strategy for the serverless inference platform, designing telemetry, dashboards, and alerts, defining SLO/SLI frameworks, and driving improvements when targets are missed
  • Building production-grade automation and tooling that reduces operational toil, improves incident response, and sets patterns that other SREs adopt
  • Owning incident management integration for inference workloads, designing frameworks, leading incident response during on-call rotations, and driving systemic improvements from post-mortems
  • Defining and implementing deployment safety practices including progressive rollouts, canary analysis, and rollback automation, establishing standards for the team
  • Partnering with product engineering teams to influence architecture decisions, ensure operational readiness, and represent the SRE perspective in design reviews
  • Mentoring Senior and mid-level SREs through code reviews, design discussions, and hands-on problem-solving

Benefits

  • Healthcare
  • 401K savings plan
  • Company holidays
  • Vacation (in the form of PTO)
  • Sick time
  • Family friendly benefits including parental leave
  • Employee assistance program with a focus on mental and financial wellness
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service