SRE

Saxon GlobalBirmingham, AL

About The Position

This role focuses on Site Reliability Engineering (SRE) principles to ensure the scalability, stability, and performance of systems. The SRE will be responsible for gathering and analyzing metrics, participating in system design and capacity planning, and working closely with the incident response team to restore services. A key aspect of the role involves balancing feature development speed with reliability and service-level objectives, while also investigating and mitigating unwanted traffic. The SRE will establish continuous process improvement cycles and partner with development teams to enhance services through testing and release procedures.

Requirements

  • 5-7 years of experience
  • Understanding of Kubernetes, containers, clusters and elastic scalability.
  • Expertise in SRE principles.
  • Mindset of continually finding ways to drive scalability, stability and performance.
  • Cloud Services experience with Google Cloud Platform (GCP).
  • Experience with API, service-based or microservice-based architecture.
  • Proficiency in infrastructure, network, database, operating systems or security troubleshooting and remediation.
  • Architecture-level knowledge of Windows and Linux and Infrastructure systems.
  • Experience with production deployment, monitoring and operational support for enterprise-class applications.
  • Experience working with Continuous Integration/ Continuous Deployment tools.
  • Experience in performance diagnostics, capacity planning, performance architecture design, performance tuning and performance monitoring.
  • Experience with Azure DevOps (ADO), Dynatrace, Prometheus, Terraform and Grafana

Nice To Haves

  • Dynatrace a plus

Responsibilities

  • Gathers and analyzes metrics from monitoring platforms to assist in performance tuning and fault tolerance.
  • Participates in system design, platform management and capacity planning.
  • Balances feature development speed and reliability with service-level objectives.
  • Works closely with the incident response team and restoring service to normal operation.
  • Understands debugging and applying troubleshooting skills.
  • Investigates, blocks and rate-limits unwanted traffic.
  • Utilizes monitoring systems and dashboards for proactive changes and alerting.
  • Establishes continuous process improvement cycles where the process, performance, and supporting technologies are reviewed and enhanced where applicable.
  • Partners with development teams to improve services through testing and release procedures.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service