Site Reliability Engineer - Mid-Level

Indotronix International CorporationSouthlake, TX
Hybrid

About The Position

Join a dynamic engineering team as a Site Reliability Engineer and play a key role in driving automation, reliability, and operational excellence across both on-premises and cloud environments. Leverage your expertise in Python, cloud infrastructure, and production operations to solve complex challenges and ensure seamless, highly available systems. This is an excellent opportunity to expand your skills in automation, observability, and AI/ML-driven operations while collaborating with forward-thinking professionals in a hybrid work environment.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or equivalent experience
  • 3–5 years of hands-on experience in Site Reliability Engineering, DevOps, Production Engineering, or related disciplines
  • Strong Python programming skills for automation and tooling
  • Practical experience managing Kubernetes clusters and supporting cloud platforms like GCP, AWS, or Azure
  • Proficiency with infrastructure automation and configuration management tools
  • Solid understanding of Linux systems, networking, and distributed application environments
  • Experience with monitoring and observability tools such as Splunk, Grafana, or Prometheus
  • Demonstrated ability to support production systems at scale and participate in incident response

Nice To Haves

  • Experience with Terraform, Ansible, or other Infrastructure as Code solutions
  • Exposure to OpenTelemetry and modern observability practices
  • Familiarity with CI/CD pipelines and deployment automation
  • Knowledge of AI/ML, AIOps, or advanced operational tooling
  • Background supporting highly available systems in regulated or enterprise environments

Responsibilities

  • Design and develop Python-based automation solutions to streamline operational processes and reduce manual effort
  • Automate and manage infrastructure across Linux, Windows, Kubernetes, GCP, and other cloud-native platforms
  • Integrate diverse tools and platforms using APIs and client libraries to create cohesive operational workflows
  • Implement infrastructure automation solutions with Ansible, Terraform, or similar tools
  • Monitor and maintain production systems to consistently meet reliability and availability targets
  • Participate in incident response, troubleshooting, and root cause analysis to resolve issues efficiently
  • Build and maintain dashboards, alerts, and monitoring systems using Splunk, Grafana, Prometheus, or comparable platforms
  • Explore and implement AI/ML-driven operational improvements, such as anomaly detection and intelligent alerting

Benefits

  • Hybrid work setup for flexibility and work-life balance
  • Opportunities for career growth and skill development in automation, cloud, and AI/ML operations
  • Collaborative and inclusive culture focused on innovation and continuous improvement
  • Work with cutting-edge technologies and modern engineering practices
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service