Observability Engineer / Site Reliability Engineer

Ontrac SolutionsChicago, IL
$90 - $100

About The Position

We are seeking an experienced Observability / Site Reliability Engineer (SRE) to design, scale, and maintain our enterprise monitoring and alerting ecosystems. In this role, you will bridge the gap between development and operations by ensuring high availability, performance tuning, and deep visibility across distributed multi-cloud and native systems. You will play a critical role in automating infrastructure and building robust observability pipelines using industry-leading cloud-native tools.

Requirements

  • Proven engineering experience within Google Cloud Platform (GCP) environments, particularly managing cloud-native monitoring and compute resources.
  • Hands-on experience with Grafana, Prometheus, and Google Cloud Observability suites.
  • Expert-level knowledge of Linux/Unix operating systems paired with strong shell scripting skills for automation and systems management.
  • Professional coding proficiency in at least one modern language (Python, Go, Java, Perl, or advanced Shell).
  • Hands-on experience managing containerized applications on Kubernetes, GKE, and/or Red Hat OpenShift.

Nice To Haves

  • Direct experience with GEM (Grafana Enterprise Metrics) is highly desirable.

Responsibilities

  • Architect, optimize, and maintain observability frameworks across cloud environments, with a specific focus on implementing Google Cloud Platform (GCP) observability tools (Cloud Logging, Cloud Monitoring, Trace, and Profiler).
  • Design, deploy, and maintain robust observability stacks across hybrid ecosystems, utilizing Prometheus, Grafana, and cloud-native integrations.
  • Drive infrastructure-as-code (IaC) initiatives using Terraform and Ansible to ensure consistent, automated deployments of infrastructure and observability tooling.
  • Build, maintain, and optimize deployment workflows within Kubernetes and Google Kubernetes Engine (GKE) / OpenShift environments using GitHub, Harness, and other CI/CD pipelines.
  • Deeply analyze Linux/Unix system administration architectures, optimizing compute resource metrics and performance tuning across complex, distributed environments.
  • Implement SRE best practices, establishing meaningful Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to ensure platform reliability.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service