Manager, Infrastructure

RZR Global Inc.San Francisco, CA
Remote

About The Position

RZR is an AI-native advertising platform built for the next era of performance marketing. We operate at the intersection of machine learning, programmatic media, and full-funnel mobile growth, powering campaigns for some of the world's most ambitious advertisers. Our platform is purpose-built to deliver outcomes at scale, not just impressions. We are a team of builders, operators, and technologists who believe the advertising industry is overdue for a fundamental rethink. We move fast, operate with a high degree of ownership, and hold ourselves to an exceptionally high standard of craft. RZR is scaling aggressively with an active M&A pipeline and a platform vision that puts us on a path to becoming an industry leader. This is a rare opportunity to join a company at an inflection point and help shape what it becomes. RZR runs its own metal — four owned-and-operated data centers across Santa Clara, Ashburn, Amsterdam, and Hong Kong, housing approximately 1,300 servers and 1.15MW of capacity, placed next to the major ad exchanges, serving 5–6M+ bid requests per second at ~20ms. This infrastructure is our competitive moat, not a cost center. As Manager, Infrastructure, you will own day-to-day and quarter-to-quarter operation of that entire footprint: the team, hardware lifecycle, capacity planning, incident response, and vendor relationships. You will take over these functions directly from the Head of Cybersecurity & Infrastructure, freeing him to focus on security and multi-entity IT. This is a hands-on manager role. There is no "scale up" button here — latency, packets-per-second, and procurement lead times are the job. The right person combines deep bare-metal operational instincts with the leadership presence to run a distributed, experienced global team from day one.

Requirements

  • 6–8 years in infrastructure or data center operations with 2+ years managing engineers
  • Bare-metal and colo depth: capacity planning, hardware procurement (Dell/Supermicro), IBX/remote-hands workflows, and the physical logistics of running owned cages
  • Network fundamentals at scale: spine-leaf architecture, BGP/peering (100G-class), transit blends, and low-latency tuning (NIC/IRQ, packets-per-second thinking)
  • Deep Linux operations: systemd, netplan, FreeIPA/Chrony, Ansible and/or Salt, Zabbix, running fleets of hundreds-plus servers
  • Experience operating large stateful distributed systems — Aerospike, Cassandra, Scylla, Kafka, or ClickHouse — under sub-50ms latency budgets and hard capacity limits
  • Demonstrated P1/P2 incident ownership: on-call program management, postmortems, and alert hygiene discipline

Nice To Haves

  • Hybrid cloud experience alongside owned metal: AWS (IAM, Route 53, GuardDuty, S3); Kubernetes exposure a plus
  • SOC 2 or compliance evidence experience; Okta and Vanta familiarity; security-minded infrastructure approach
  • Experience in adtech, RTB, or other high-QPS, latency-sensitive environments
  • Netris or other SDN controller experience; MAAS provisioning familiarity

Responsibilities

  • Manage and develop the InfraOps team across US and APAC time zones, including regional DC owners (SV/VA and NL/HK) and network engineering
  • Own the weekly DevOps check-in cadence, alert reviews, and 24/7 on-call coverage model
  • Drive P1/P2 incident response end to end — accountability for MTTR reduction, runbook coverage, and alert hygiene
  • Own capacity planning and hardware lifecycle across all four data centers: Dell and Supermicro procurement through VARs, GPU expansion for on-prem ML training and inference, colo power and space management, and remote-hands logistics with Equinix and Digital Realty
  • Run the annual cloud-vs-colo evaluation alongside leadership, with full ownership of the recommendation
  • Oversee the core infrastructure stack: Ubuntu/systemd fleet, FreeIPA, Ansible/Salt configuration management, MAAS provisioning, and Zabbix monitoring
  • Manage the spine-leaf Mellanox/NVIDIA network via Netris, including 100G Google peering and transit blend (Lumen/Cogent/Zayo)
  • Support the stateful data tier — Aerospike, Kafka, ClickHouse, Hadoop/HDFS — across capacity limits, evictions, migrations, and low-latency tuning
  • Own colo and vendor relationships and budgets: Equinix and Digital Realty invoices, transit contracts, VAR procurement, and Netris licensing
  • Partner with the security/compliance function on SOC 2 Type 2 evidence, infrastructure hardening, and access reviews
  • Operate comfortably within a multi-entity environment (RZR/Skillz/Firy/Beamable shared IT) with comfort in M&A-flavored ambiguity

Benefits

  • Periodic travel to data center sites (Santa Clara, CA; Ashburn, VA; Amsterdam; Hong Kong) and SF HQ for operational reviews, team time, and site work.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service