Site Reliability Engineer II

Akamai
1d$95,000 - $171,000Remote

About The Position

The Akamai Inference Cloud team is part of Akamai's Cloud Technology Group. We design, implement, deploy and operate AI platforms that enable customers to run inference models and developers to create AI applications. In this role, responsibilities will include automation, monitoring, incident response, and working collaboratively with skilled team members. Candidates should possess expertise in Linux systems, automation, and SRE practices. Daily activities involve coding, improving dashboards, enhancing alerts, and minimizing repetitive tasks. Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform.

Requirements

  • Have 2+ years of experience in Site Reliability Engineering and a Bachelor's Degree or its equivalent experience
  • Demonstrate coding ability in at least one programming language (Python or Go) with experience writing automation
  • Have experience with Linux systems administration and the ability to troubleshoot complex infrastructure issues
  • Show familiarity with Kubernetes and containerization concepts
  • Have experience with monitoring and observability tools such as Prometheus, Grafana, or similar
  • Have exposure to CI/CD pipelines and infrastructure-as-code tools (Terraform, SaltStack, or equivalent)
  • Show a willingness to learn and grow, with genuine curiosity about AI infrastructure and distributed systems

Responsibilities

  • Building and maintaining dashboards, alerts, and monitoring for inference workloads using Akamai's existing observability platform
  • Writing automation and tooling in Python or Go to reduce operational toil and improve system reliability
  • Building and improving runbooks for inference-specific operational procedures, integrating into Akamai's existing incident management processes
  • Contributing to SLO tracking and reporting, identifying trends and areas for improvement
  • Supporting CI/CD pipeline maintenance, deployment safety checks, and rollback procedures
  • Collaborating with product engineering teams to troubleshoot complex problems across the stack
  • Participating in on-call rotations, responding to production incidents, and conducting blameless post-mortems

Benefits

  • At Akamai, we will provide you with opportunities to grow, flourish, and achieve great things. Our benefit options are designed to meet your individual needs for today and in the future. We provide benefits surrounding all aspects of your life: Your health Your finances Your family Your time at work Your time pursuing other endeavors
  • Akamai provides industry-leading benefits including healthcare, 401K savings plan, company holidays, vacation (in the form of PTO), sick time, family friendly benefits including parental leave and an employee assistance program including a focus on mental and financial wellness; Eligibility requirements apply.
© 2024 Teal Labs, Inc
Privacy PolicyTerms of Service