Senior Software Engineer, Platform & Infrastructure - Riot Technology

Riot GamesLos Angeles, CA
$161,500 - $227,000

About The Position

Platform and infrastructure engineers at Riot build the foundational systems that enable teams to develop, deploy, and operate systems at global scale. They partner across disciplines with other engineers, data scientists, designers, and product teams to ensure reliable, scalable, and secure infrastructure underpins every ML capability that reaches players. As a Senior Platform & Infrastructure Engineer on the Riot Technology team, you will design, build, and operate the core infrastructure and ML platforms. Your focus will be on the computing and orchestration platforms that power large-scale distributed training of agents (e.g. via RL, IL, and other techniques), simulation environments, and policy evaluation, as well as the CI/CD, infrastructure-as-code, observability, and developer tooling that keep these systems production-grade. You will close critical infrastructure gaps across the team's stack, driving improvements to standards, automation, and operational maturity. You will operate independently on multi-month work efforts and begin to influence technical direction beyond your immediate team. You will report to the Manager of Machine Learning.

Requirements

  • Bachelor’s degree in Computer Science or a related field, or equivalent practical experience.
  • 3+ years of software engineering experience, with meaningful experience in infrastructure, platform engineering, or SRE roles.
  • Experience operating distributed systems in production and keeping them healthy under real load.
  • Strong experience with Kubernetes, AWS or GCP, infrastructure-as-code, CI/CD, deployment automation, and production tooling.
  • Experience with GPU compute infrastructure, including scheduling, multi-node orchestration, and resource optimization for long-running training workloads.
  • Proficiency in Python and solid understanding of networking, microservices, core infrastructure services, and distributed systems fundamentals.

Nice To Haves

  • Familiarity with MLOps workflows such as model versioning, pipeline orchestration, experiment tracking, artifact management, and reproducible ML workflows.
  • Experience with distributed training or HPC frameworks, inference serving, systems languages, high-performance networking, Unreal/client-server architecture, AI-assisted development tools
  • Passion for games and player experience.

Responsibilities

  • Build and operate Kubernetes, multi-node GPU clusters, and networking infrastructure for distributed ML bot training and large-scale policy evaluation.
  • Design infrastructure for running simulation environments at scale, enabling parallel rollouts, data collection, training, and evaluation.
  • Build CI/CD, deployment automation, artifact management, and infrastructure-as-code across cloud environments.
  • Improve platform reliability, cost efficiency, performance, reproducibility, auditability, and operational maturity.
  • Build observability, monitoring, alerting, health indicators, and SLO-aligned dashboards for infrastructure and ML workloads.
  • Develop internal APIs, control planes, templates, and developer tooling for distributed training and evaluation workflows.
  • Support MLOps workflows including automated training pipelines, model artifact management, experiment tracking, and reproducible ML lifecycle operations.
  • Build security and governance controls, manage production incidents, drive root-cause remediation, mentor engineers, and support recruiting for platform roles.

Benefits

  • open paid time off policy
  • flexible work schedules
  • medical, dental, and life insurance
  • parental leave
  • 401k with company match
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service