About The Position

We are looking for skilled Senior Site Reliability Engineers with a minimum of 5 years of experience to join a dynamic team within a leading organization. This role involves supporting and improving cloud operations for microservice-based platforms, with a focus on production reliability, incident response, cloud infrastructure, automation, observability, Kubernetes operations, and CI/CD workflows across AWS and Azure environments.

Requirements

  • Over 5 years of experience as a Senior Site Reliability Engineer.
  • Strong hands-on experience with AWS and Azure cloud platforms.
  • Strong experience with Terraform for Infrastructure as Code (IaC).
  • Experience with Atlantis, ArgoCD, or similar infrastructure and deployment automation tools.
  • Strong hands-on experience with Docker and Kubernetes.
  • Experience designing, maintaining, and troubleshooting complex CI/CD pipelines.
  • Strong production support experience, including incident management, Root Cause Analysis (RCA), postmortems, and runbook creation.
  • Strong observability experience, including monitoring, alerting, logging, diagnostics, and performance analysis.
  • Good understanding of cloud networking, security, access controls, and InfoSec practices.
  • Experience with version control, branching, merging, pull requests, and conflict resolution.
  • Understanding of cloud cost optimization and resource utilization.
  • Bachelor’s degree or higher.
  • Fluent in English (Advanced).
  • Excellent communication, empathy, commitment, leadership, teamwork, and a proactive attitude.
  • Only residents of Mexico.

Nice To Haves

  • Experience with microservice-based platforms.
  • Experience with Datadog, CloudWatch, Grafana, Prometheus, Splunk, AppDynamics, or similar tools.
  • Scripting or programming experience using Python, Bash, Go, or Java.
  • Experience with SLI/SLO/SLA, error budgets, capacity planning, and resilience engineering.
  • Experience with disaster recovery testing and production readiness reviews.
  • Prior experience mentoring junior engineers or leading technical troubleshooting.

Responsibilities

  • Own and improve the reliability of cloud-based services and supporting infrastructure.
  • Participate in on-call rotations and support production systems outside normal business hours.
  • Lead incident response activities, including triage, escalation, mitigation, and service restoration.
  • Drive blameless postmortems and ensure corrective actions are tracked to closure.
  • Design, implement, and maintain Infrastructure as Code using Terraform and tools such as Atlantis.
  • Manage and enhance GitOps and deployment workflows using ArgoCD and related CI/CD tools.
  • Support and improve cloud and container platforms across AWS and Azure.
  • Manage Kubernetes-based workloads, containers, virtual servers, and distributed systems.
  • Build automation to reduce manual effort and improve operational efficiency.
  • Configure and improve monitoring, alerting, logging, diagnostics, and observability.

Benefits

  • Attractive Salary
  • Premium Benefits
  • Performance bonuses
  • Grocery coupons
  • Savings
  • Aguinaldo
  • Premium vacations
  • Vacations paid
  • SGMM Medical insurance
  • Family insurance
  • Life insurance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service