Site Reliability Engineer (SRE)

Smart IMSSouthlake, TX
Onsite

About The Position

As a Site Reliability Engineer (SRE), you will be responsible for improving the reliability, scalability, and operational efficiency of production systems through automation, observability, and incident management. The ideal candidate will have strong Python development expertise, hands-on experience supporting large-scale production environments, and a proven track record of reducing operational toil through automation.

Requirements

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent professional experience.
  • 3-5 years of hands-on Site Reliability Engineering or Production Engineering experience supporting large-scale production systems.
  • Strong Python programming expertise with demonstrated experience building automation tools, frameworks, and operational solutions.
  • Proven experience in production operations, incident response, root cause analysis, and reliability engineering.
  • Experience automating operational processes and reducing manual toil through engineering solutions.
  • Hands-on experience with Kubernetes and cloud platforms such as GCP, AWS, or Azure.
  • Experience with monitoring and observability tools including Splunk, Grafana, Prometheus, Datadog, or similar platforms.
  • Strong understanding of Linux systems, networking concepts, and distributed application architectures.
  • Experience with Infrastructure as Code and configuration management tools such as Terraform and Ansible.
  • Excellent analytical, troubleshooting, and problem-solving skills with the ability to perform effectively in mission-critical environments.

Responsibilities

  • Develop Python-based automation solutions to eliminate manual operational tasks and improve efficiency.
  • Support and maintain large-scale production systems, ensuring high availability and reliability.
  • Participate in incident response, troubleshooting, root cause analysis, and problem remediation activities.
  • Automate infrastructure management across cloud, Kubernetes, Linux, and Windows environments.
  • Implement and support infrastructure automation using tools such as Terraform and Ansible.
  • Build and maintain observability solutions including dashboards, alerts, metrics, logs, and monitoring frameworks.
  • Drive operational improvements to reduce recurring incidents and improve system stability.
  • Perform performance analysis, capacity planning, and system health assessments.
  • Support disaster recovery, failover testing, and operational readiness initiatives.
  • Evaluate and implement emerging observability, automation, and AIOps capabilities.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service