About The Position

We are seeking a Senior Site Reliability Engineer (SRE) with expertise in Azure DevOps and Kubernetes. This role involves supporting large-scale distributed applications and production environments, focusing on observability, monitoring, CI/CD pipelines, automation, and incident management. The ideal candidate will have a strong background in cloud platforms and container technologies, along with excellent troubleshooting and performance tuning skills.

Requirements

  • 5 years of hands-on experience in Site Reliability Engineering (SRE).
  • Strong expertise in observability and monitoring tools such as Splunk, Datadog, Dynatrace, Grafana, Prometheus, or OpenTelemetry.
  • Experience in telemetry data collection, logging, metrics, tracing, and dashboard development.
  • Proficiency with cloud platforms (AWS, Azure, or GCP) and container technologies (Kubernetes, Docker).
  • Hands-on experience with CI/CD pipelines, automation, and infrastructure as code.
  • Strong troubleshooting, performance tuning, and incident management skills.
  • Knowledge of scripting/programming languages such as Python, Bash, or Go.
  • Experience supporting large-scale distributed applications and production environments.
  • 8-10 years of minimum experience.

Nice To Haves

  • Microsoft FABRIC

Responsibilities

  • Implement and manage observability and monitoring tools (Splunk, Datadog, Dynatrace, Grafana, Prometheus, OpenTelemetry).
  • Develop telemetry data collection, logging, metrics, tracing, and dashboard solutions.
  • Support and manage cloud platforms (AWS, Azure, or GCP) and container technologies (Kubernetes, Docker).
  • Build and maintain CI/CD pipelines, automation scripts, and infrastructure as code.
  • Troubleshoot, tune performance, and manage incidents in production environments.
  • Write scripts or programs in languages such as Python, Bash, or Go.
  • Support large-scale distributed applications and production environments.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service