Head of Platform & Cloud Engineering

RapidRatingsQuincy, MA
Hybrid

About The Position

RapidRatings is a financial technology company focused on enhancing transparency and resilience in the global economy through quantitative financial health intelligence. This role is for a deeply technical senior leader responsible for the cloud platform, resilience, security, and AI infrastructure that powers RapidRatings' products. The position combines deep systems engineering with modern SRE practices, covering low-level cloud architecture, high-scale distributed systems, and platform engineering. The Head of Platform & Cloud Engineering will lead the SRE team, establish operational excellence, and build the foundational infrastructure for AI systems, agents, and product squads, leveraging AI and developer self-service to reduce routine tasks.

Requirements

  • Deep, hands-on experience running production Kubernetes (EKS) at scale.
  • Deep, hands-on experience across the native AWS stack (EKS/ECS, IAM, VPC networking, CloudFront, Lambda, RDS/DynamoDB).
  • Proficiency with infrastructure as code (Terraform / OpenTofu).
  • Strong foundation in Linux internals, networking (TCP/IP, DNS, routing), performance tuning, and operating distributed systems at scale.
  • Direct experience implementing and maintaining SOC 2 and ISO 27001 technical controls within cloud-native environments.
  • Proficiency in Python, Go, or Bash for building platform tooling, custom telemetry, and automated remediation.
  • Hands-on exposure to model-hosting patterns, vector databases, API gateways, and LLM orchestration infrastructure.
  • Track record of managing distributed SRE/infrastructure teams, including international squads.

Nice To Haves

  • Treat AI as a strong collaborator, building platforms and guardrails that lift the whole engineering group.
  • Bring a calm, positive attitude under pressure, including during production incidents and against tight deadlines.

Responsibilities

  • Architect high-availability, multi-region AWS infrastructure optimized for scale, latency, resilience, and structural reliability.
  • Lead the SRE function to establish robust telemetry, SLO frameworks, and rigorous root-cause analysis through blameless post-incident reviews.
  • Drive self-healing automation that minimizes operational overhead and mean-time-to-recovery.
  • Build and operate the model-serving layer, routing pipelines, cost attribution, and governance frameworks for AI agents, copilots, and Platform Model Context Protocols (MCPs).
  • Own the platform's threat detection and defense, embedding SOC 2 Type II and ISO 27001 controls into infrastructure code and automated guardrails.
  • Design self-service infrastructure patterns for engineering squads to ship securely and efficiently, while continuously driving cloud cost optimization.
  • Manage distributed SRE/infrastructure teams, including international squads, while remaining directly involved in architectural design.
  • Set technical direction and influence others through standards, design review, documentation, and mentoring.
  • Communicate clearly across all levels, from technical specialists to the executive group.

Benefits

  • Bonus
  • Flexible work environment
  • Unlimited PTO
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service