Senior Site Reliability Engineer, Infrastructure

DraftKings Inc.Boston, MA
$128,000 - $160,000Onsite

About The Position

At DraftKings, AI is becoming an integral part of both our present and future, powering how work gets done today, guiding smarter decisions, and sparking bold ideas. It’s transforming how we enhance customer experiences, streamline operations, and unlock new possibilities. Our teams are energized by innovation and readily embrace emerging technology. We’re not waiting for the future to arrive. We’re shaping it, one bold step at a time. To those who see AI as a driver of progress, come build the future together. The Crown Is Yours As a Senior Site Reliability Engineer, you’ll build and scale the critical Kubernetes infrastructure that powers our platforms and services. You’ll solve complex reliability challenges across public cloud and on-premise environments, designing automation-first solutions that strengthen performance and simplify operations. You’ll help shape architectural decisions, advance stability at scale, and build tools that give our teams the confidence to move quickly and deliver reliably.

Requirements

  • A Bachelor’s Degree in Computer Science or a related field, or equivalent education, experience, and training.
  • At least 4 years of experience managing distributed cloud and on-premise environments at scale, including strong hands-on experience with Amazon Web Services; experience with Google Cloud Platform, vSphere, or Nutanix is a plus.
  • Deep expertise in Kubernetes and container orchestration, with experience designing, scaling, and troubleshooting complex workloads.
  • Strong software development experience using languages such as Go and Python to build automation and infrastructure tooling.
  • Working knowledge of networking and Linux-based systems, including container runtimes such as Docker and containerd, packet-level debugging, and kernel troubleshooting.
  • Experience with Infrastructure as Code and configuration management tools to build scalable, consistent, and repeatable infrastructure.

Nice To Haves

  • experience with Google Cloud Platform, vSphere, or Nutanix is a plus.

Responsibilities

  • Drive stability, performance, and scalability across our global compute platform spanning multiple public clouds and on-premise environments.
  • Build self-healing, fault-tolerant infrastructure and internal tooling that automates repetitive operational work and reduces toil for Platform and Application teams.
  • Operate and evolve our GitOps delivery model, using Rancher Fleet, Flux, and Helm to deploy core Kubernetes services and application workloads consistently and reliably.
  • Own Kubernetes scaling and capacity strategies using technologies including Karpenter, Horizontal Pod Autoscaler (HPA), Kubernetes Event-Driven Autoscaling (KEDA), and predictive scaling based on event and calendar data.
  • Define and monitor service-level objectives and reliability metrics for platform components using Datadog and our logging pipeline.
  • Strengthen our engineering practices by sharing knowledge, contributing to architectural and design discussions, and participating in an on-call rotation.

Benefits

  • bonus
  • equity
  • benefits as applicable
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service