Head of Platform & Cloud Engineering

RapidRatingsQuincy, MA
$145,000 - $174,500Hybrid

About The Position

We are hiring a deeply technical senior leader to own the cloud platform, resilience, security posture, and AI infrastructure powering RapidRatings' products. The role blends deep systems engineering with modern SRE practice, bridging low-level cloud architecture, high-scale distributed systems, and platform engineering. As AI and developer self-service absorb routine pipeline toil, you will raise the technical bar for operational excellence, lead our SRE team, and build the paved-road infrastructure our AI systems, agents, and product squads run on.

Requirements

  • Deep, hands-on experience running production Kubernetes (EKS) at scale.
  • Deep, hands-on experience across the native AWS stack (EKS/ECS, IAM, VPC networking, CloudFront, Lambda, RDS/DynamoDB) and infrastructure as code (Terraform / OpenTofu).
  • Strong foundation in Linux internals, networking (TCP/IP, DNS, routing), performance tuning, and operating distributed systems at scale.
  • Direct experience implementing and maintaining SOC 2 and ISO 27001 technical controls within cloud-native environments.
  • Proficiency in Python, Go, or Bash for building platform tooling, custom telemetry, and automated remediation.
  • Hands-on exposure to model-hosting patterns, vector databases, API gateways, and LLM orchestration infrastructure.
  • Track record of managing distributed SRE/infrastructure teams, including international squads, while remaining directly involved in architectural design.

Nice To Haves

  • A natural problem solver who stays curious, works logically, and digs past symptoms to root cause.
  • Treats AI as a strong collaborator, building the platforms and guardrails that lift the whole engineering group rather than only your own output.

Responsibilities

  • Architect high-availability, multi-region AWS infrastructure optimized for scale, latency, resilience, and structural reliability.
  • Lead the SRE function to establish robust telemetry, SLO frameworks, and rigorous root-cause analysis through blameless post-incident reviews.
  • Drive self-healing automation that minimises operational overhead and mean-time-to-recovery.
  • Build and operate the model-serving layer, routing pipelines, cost attribution, and governance frameworks for our AI agents, copilots, and Platform Model Context Protocols (MCPs).
  • Own the platform's threat detection and defense.
  • Embed SOC 2 Type II and ISO 27001 controls directly into infrastructure code and automated guardrails.
  • Design self-service infrastructure patterns so engineering squads ship securely and without friction, while continuously driving cloud cost optimization.
  • Set technical direction and carry others with you through standards, design review, documentation, and mentoring.
  • Communicate clearly across every tier, from technical specialists to the executive group, with meticulous attention to detail.
  • Bring a calm, positive attitude under pressure, including during production incidents and against tight deadlines.

Benefits

  • Bonus
  • Flexible work environment
  • Unlimited PTO
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service