Head of Platform & Cloud Engineering

RapidRatingsNew York, NY
Hybrid

About The Position

RapidRatings is seeking a deeply technical senior leader to oversee the cloud platform, resilience, security posture, and AI infrastructure that powers its products. This role combines deep systems engineering with modern SRE practices, encompassing low-level cloud architecture, high-scale distributed systems, and platform engineering. The Head of Platform & Cloud Engineering will elevate the technical standards for operational excellence, lead the SRE team, and develop the foundational infrastructure for AI systems, agents, and product squads, while AI and developer self-service handle routine tasks.

Requirements

  • Deep, hands-on experience running production Kubernetes (EKS) at scale.
  • Deep, hands-on experience across the native AWS stack (EKS/ECS, IAM, VPC networking, CloudFront, Lambda, RDS/DynamoDB) and infrastructure as code (Terraform / OpenTofu).
  • Strong foundation in Linux internals, networking (TCP/IP, DNS, routing), performance tuning, and operating distributed systems at scale.
  • Direct experience implementing and maintaining SOC 2 and ISO 27001 technical controls within cloud-native environments.
  • Proficiency in Python, Go, or Bash for building platform tooling, custom telemetry, and automated remediation.
  • Hands-on exposure to model-hosting patterns, vector databases, API gateways, and LLM orchestration infrastructure.
  • Track record of managing distributed SRE/infrastructure teams, including international squads, while remaining directly involved in architectural design.
  • Ability to set technical direction and influence others through standards, design review, documentation, and mentoring.
  • Clear communication skills across technical specialists and executive groups.
  • Calm, positive attitude under pressure, including during production incidents and against tight deadlines.

Nice To Haves

  • Treat AI as a strong collaborator, building the platforms and guardrails that lift the whole engineering group rather than only your own output.

Responsibilities

  • Architect high-availability, multi-region AWS infrastructure optimized for scale, latency, resilience, and structural reliability.
  • Lead the SRE function to establish robust telemetry, SLO frameworks, and rigorous root-cause analysis through blameless post-incident reviews.
  • Drive self-healing automation that minimizes operational overhead and mean-time-to-recovery.
  • Build and operate the model-serving layer, routing pipelines, cost attribution, and governance frameworks for AI agents, copilots, and Platform Model Context Protocols (MCPs).
  • Own the platform's threat detection and defense, embedding SOC 2 Type II and ISO 27001 controls directly into infrastructure code and automated guardrails.
  • Design self-service infrastructure patterns for engineering squads to ship securely and without friction, while continuously driving cloud cost optimization.
  • Directly manage the SRE team and set cloud and platform standards across all engineering squads.

Benefits

  • Bonus
  • Flexible work environment
  • Unlimited PTO
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service