Senior/Staff Site Reliability Engineer - Data Center

PathAIBoston, MA
$146,250 - $225,000Hybrid

About The Position

PathAI's mission is to improve patient outcomes with AI-powered pathology. PathAI is transforming traditional pathology methods into powerful, new technologies. These innovations in pathology can help accelerate drug development, improve confidence in the accuracy of diagnosis, and get life-saving therapies to patients more quickly. At PathAI, you'll work with a diverse and talented team of people, who are dedicated to solving complex problems and making a huge impact. We are expanding our team and recruiting for a skilled Senior/Staff Site Reliability Engineer focused on designing, building, and operating our on-prem/cloud environment.

Requirements

  • BS in Computer Science, Computer Engineering, Electrical Engineering, Software Engineering or closely related technical field.
  • 8 years experience working in physical hardware/facilities, networking, automation or other relevant areas.
  • Demonstrated experience with modern datacenter network designs and comfort operating across network layers.
  • Administered physical hardware stacks in production settings (iDRAC/IPMI/Nvidia UFM/Juniper Systems).
  • Demonstrated experience and opinions on virtualization, containerization, or container orchestration platforms. (EKS-Anywhere/ClusterAPI/KVM).
  • Strong expertise in storage solutions and optimizing them for high-performance workloads (e.g., Quobyte, S3, FSx, EFS).
  • Highly proficient with automation tools; eliminate toil by automating everything through scripting, configuration management tools (Ansible/RedFish).
  • Built monitoring infrastructure with modern observability tools (Datadog/Grafana/Prometheus).
  • Proven operational background managing critical production systems, with extensive experience in incident response, infrastructure scaling, and navigating high-growth challenges.
  • Ability to travel to onsite Datacenter location(s) as needed.

Nice To Haves

  • Outstanding interpersonal, verbal, and written communication and influencing skills: have built and cultivated important relationships both inside and outside of the organization and externally; have proven abilities to influence internal partners and stakeholders, thought leaders, national advocacy organizations, national standard-setting bodies, and other relevant external parties.
  • Strong analytical and critical thinking skills with attention to detail; ability to manage multiple projects and drive results in a fast-paced environment; collaborative mindset with demonstrated leadership capabilities.

Responsibilities

  • Advance the state of our operations by implementing SRE best practices - focusing on users, monitoring, and automation.
  • Design, build and operate our data center to support our rapidly growing Machine Learning team.
  • Build highly-secure on-premises environments handling NIST/ISO standards.
  • Integrate on-premises datacenter environments with existing cloud infrastructure to create a seamless hybrid cloud environment.
  • Improve the reliability and resilience of our infrastructure through root-cause analysis and reviewing gaps in designs, and implementations of our infrastructure.
  • Participate in platform on-call rotations and assist with urgent incident response.

Benefits

  • Relocation benefits are not available for this position.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service