Site Reliability Engineer II

AkamaiCambridge, MA
CRC 16,932,850 - CRC 30,479,150Hybrid

About The Position

Our team designs, develops, and manages applications and infrastructure that support Akamai's Compute products and services. We do this while maintaining Akamai's mission at the forefront of what we do. Make life better for billions of people, billions of times a day. In this role, you'll create solutions to improve automation and efficiency for systems and teams. Responsibilities include optimizing workflows, infrastructure, and applications. You'll collaborate on deployment, monitoring, and resolving incidents. You'll focus on reliability, scalability, and efficiency through automation and resource optimization. You'll promote continuous improvement and operational excellence across all systems.

Requirements

  • 2 years of relevant experience and a Bachelor's degree in Computer Science or its equivalent
  • Possess experience in a SysAdmin (Linux/Unix Administration), DevOps or SRE role, working with large scale distributed systems
  • Demonstrate experience in Kubernetes and large-scale containerization systems.
  • Possess at least one programming language (Python/Golang) and configuration management with Terraform/SaltStack/Ansible
  • Define SLOs and work with observability tools like Prometheus, Grafana, and distributed tracing to enhance system monitoring.
  • Demonstrate accountability for reliability, develop automation and monitoring, and collaborate effectively with an engineering team unfamiliar with SRE practices.

Responsibilities

  • Providing support and mentorship for other engineers within the department
  • Developing and maintaining automated tools and scripts to enhance system reliability, deployment processes, and incident response efficiency.
  • Improving our system monitoring to speed error detection and remediation, enhancing performance and reliability of virtualization platform
  • Participating in on-call rotations, guiding restoration and repair of service-impacting issues
  • Writing automation and tooling to reduce operational toil, improve deployment safety, and accelerate incident response
  • Contributing to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructure

Benefits

  • healthcare
  • 401K savings plan
  • company holidays
  • vacation (in the form of PTO)
  • sick time
  • family friendly benefits including parental leave
  • an employee assistance program including a focus on mental and financial wellness
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service