Senior Site Reliability Engineer (SRE) - Hybrid

Smart IMSAustin, TX
Hybrid

About The Position

As a Senior Site Reliability Engineer (SRE), you will lead reliability, observability, automation, and operational excellence initiatives for enterprise-scale applications and platforms. This role combines software engineering, infrastructure operations, and AI/ML-driven observability to enhance system reliability, reduce operational toil, and drive automation across mission-critical environments.

Requirements

  • Bachelor's degree in Computer Science, Engineering, or a related technical field.
  • 6-8 years of experience supporting and administering enterprise-scale production environments.
  • 6-8 years of experience developing automation scripts, monitoring solutions, dashboards, and alerting frameworks.
  • Strong experience with Linux and Windows system administration, troubleshooting, tuning, and deployments.
  • Proficiency in one or more programming languages including Python, PowerShell, Java, .NET, or Bash.
  • Experience with cloud platforms, application deployments, migrations, and high-availability architectures.
  • Knowledge of networking concepts including DNS, DHCP, firewalls, routing, and distributed systems.
  • Experience with observability tools such as Splunk, AppDynamics, or similar monitoring platforms.
  • Strong understanding of databases and messaging technologies including SQL, Oracle, MongoDB, Kafka, RabbitMQ, Solace, or IBM MQ.
  • Proven experience applying AI/ML, AIOps, predictive alerting, or ML-driven observability solutions in production environments.

Responsibilities

  • Design and implement automation solutions that improve operational efficiency and platform reliability.
  • Lead AI/ML-driven observability, anomaly detection, predictive alerting, and AIOps initiatives.
  • Develop scripts, tools, and frameworks to automate infrastructure and operational processes.
  • Expand automation coverage across deployment, monitoring, alerting, and self-healing workflows.
  • Collaborate with engineering, operations, and product teams to improve system availability and performance.
  • Troubleshoot critical production issues and drive root cause analysis and remediation efforts.
  • Build and maintain monitoring dashboards, alerting frameworks, and operational intelligence solutions.
  • Support CI/CD, GitOps, and deployment automation strategies to accelerate software delivery.
  • Conduct capacity planning, performance analysis, and operational readiness assessments.
  • Participate in on-call support and ensure reliable operation of mission-critical systems.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service