Principal Site Reliability Engineer - Paze

Early Warning®Chicago, IL
$194,000 - $284,000Hybrid

About The Position

The Principal Site Reliability Engineer applies software engineering and systems engineering practices to improve the reliability, resilience, scalability, and operational health of production services. The role partners with Software Engineering and other technology teams to ensure reliability, observability, recoverability, performance, and operational readiness are engineered into systems throughout their lifecycle. The role operates at enterprise scope, establishing technical direction and applying evidence-driven engineering, technical rigor, sound judgment, automation, and broad systems expertise across organizational boundaries.

Requirements

  • Typically 15+ years of relevant professional experience in Software Engineering, Site Reliability Engineering, Systems Engineering, Cloud/Platform Engineering, DevOps, Infrastructure Engineering, Architecture where applicable, or a comparable technical discipline.
  • Experience with software development or scripting using one or more modern programming languages.
  • Experience with software engineering principles, distributed systems, production troubleshooting, automation, and observability appropriate to the level.
  • Experience with public cloud technologies and architectures, preferably AWS, along with infrastructure, networking, Linux/Unix, and modern application architectures appropriate to the level.
  • Demonstrated analytical, problem-solving, communication, and collaboration skills appropriate to the scope of the role.
  • Candidates responding to this posting must independently possess the eligibility to work in the United States, for any employer, at the date of hire. This position is ineligible for employment Visa sponsorship.

Nice To Haves

  • Hands-on experience with AWS is preferred, or comparable experience with another major cloud platform such as Microsoft Azure, Google Cloud Platform (GCP), or Oracle Cloud Infrastructure (OCI).
  • Experience developing, deploying, operating, or improving highly available production software or distributed systems.
  • Experience with CI/CD, Infrastructure as Code, containers or orchestration, observability, monitoring, alerting, and software-delivery automation.
  • Experience with SLIs, SLOs, error budgets, incident management, performance analysis, capacity management, resilience testing, disaster recovery, or operational readiness appropriate to the level.
  • Experience creating reusable automation, tooling, platforms, patterns, or practices that improve engineering effectiveness.
  • Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, Information Systems, or a related technical field, or equivalent practical experience.

Responsibilities

  • Use software engineering, automation, and DevOps principles and practices to continually improve how services are built, tested, deployed, observed, operated, and recovered.
  • Use data, evidence, experimentation, and rigorous engineering analysis appropriate to the level to identify reliability risks, test assumptions, and guide technical decisions.
  • Define, implement, or improve SLIs, SLOs, error budgets, and other service-health measures appropriate to the scope of responsibility.
  • Improve observability through metrics, logging, tracing, monitoring, alerting, dashboards, and service-health instrumentation.
  • Drive continuous improvement across CI/CD, observability, deployment practices, Infrastructure as Code, automation, testing, incident response, capacity management, resilience, and operational readiness.
  • Identify recurring or systemic production issues and translate operational experience into improvements in code, architecture, automation, tooling, and engineering practices.
  • Partner with Software Engineering teams to incorporate reliability, resiliency, scalability, performance, observability, recoverability, and operational readiness throughout the development lifecycle.
  • Participate in or lead incident response, troubleshooting, service restoration, and blameless post-incident learning appropriate to the level.
  • Provides enterprise-level technical leadership for critical production incidents and establishes or influences engineering practices that improve incident response, escalation, service restoration and sustainable on-call operations across the organization.
  • Reduce operational toil and unnecessary manual intervention through software, automation, reusable patterns, and better engineering practices.

Benefits

  • Competitive medical (PPO/HDHP), dental, and vision plans as well as company contributions to your Health Savings Account (HSA) or pre-tax savings through flexible spending accounts (FSA) for commuting, health & dependent care expenses.
  • 100% Company Safe Harbor Match on your first 6% deferral immediately upon eligibility.
  • Flexible Time Off for Exempt (salaried) employees, as well as generous PTO for Non-Exempt (hourly) employees, plus 11 paid company holidays and a paid volunteer day.
  • 12 weeks of Paid Parental Leave
  • Maven Family Planning – provides support through your Parenting journey including egg freezing, fertility, adoption, surrogacy, pregnancy, postpartum, early pediatrics, and returning to work.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service