Principal Site Reliability Engineer, Platform

Blue River Technology
$174,000 - $305,000Remote

About The Position

We are seeking a Principal Site Reliability Engineer to join the Platform organization, which accelerates company-wide adoption and scaling of automation and robotics. The Platform's product is a set of API services and infrastructure designed to overcome scaling hurdles, such as operational complexity and system exceptions, thereby enabling the rapid launch and scaling of new autonomy innovations and products at Blue River. In this role, you will join a fun, fast-moving engineering team to drive architectural decisions, mentor engineers across teams, and shape our platform's direction.

Requirements

  • Min. of 8 years of deep experience in building and maintaining infrastructure for data-intensive, high-availability applications, including six years building and maintaining public cloud solutions.
  • Deep understanding of cloud orchestration tools such as Kubernetes and Terraform.
  • Deep understanding of software design methodologies, information systems architecture, object-oriented design, and software design patterns.
  • Deep understanding of securing cloud infrastructure (preferably AWS and Kubernetes).
  • Deep experience in one or more of the following languages: Golang (preference), Python, JavaScript, Rust.
  • Deep experience in CI/CD tooling (GitHub Actions, ArgoCD, ArgoCD Image Updater, Artifactory).

Nice To Haves

  • You are interested in robotic applications and developing software that assists robots.
  • You are excited about robotics and the future of automation.
  • You are a self-starter with infectious enthusiasm, energy, and problem-solving abilities.

Responsibilities

  • Architect, scale, and own essential infrastructure.
  • Build and maintain a Kubernetes-based platform supporting multiple teams and services.
  • Build backend services (Golang) to support autonomous systems.
  • Partner with product teams to launch new products on the platform.
  • Grow our high availability infrastructure while maintaining key metrics such as uptime.
  • Build tooling to support our platform and development teams.
  • Perform end-to-end performance analysis, identify areas for improvement, and implement robust solutions.
  • Work with cloud vendors and external technical support for upgrades and rapid problem resolution.
  • Participate in on-call rotation, triaging and resolving production incidents with thorough root cause analysis and postmortem documentation.
  • Design and maintain observability infrastructure, dashboards, alerts, and log aggregation to ensure visibility into platform health and service performance.
  • Collaborate with the security team to conduct regular risk assessments.
  • Maintain the risk register and develop and implement mitigation plans.
  • Assess intrusion detection alerts. Improve systems and services that digest threat feeds.
  • In collaboration with IT and purchasing teams, establish and maintain the payment process for each SaaS service.

Benefits

  • Annual performance bonus
  • Competitive benefit package
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service