Application Support SRE

S&P Global MobilitySouthfield, MI

About The Position

This role focuses on Production Support & Reliability, providing hands-on support for enterprise applications in production. The SRE will troubleshoot issues across application and infrastructure layers, drive incident triage and resolution, and ensure systems remain stable, available, and supportable. The position involves participating in a 24/7 on-call rotation, responding to and resolving production incidents, and coordinating with engineering teams during high-severity issues. Responsibilities also include executing deployments using Ansible and CI/CD pipelines, using Git workflows for release management, supporting deployment validation and troubleshooting, and improving deployment reliability. The role requires monitoring systems using tools like Splunk, Google Analytics, CloudWatch, and synthetic monitoring tools, analyzing logs, metrics, and alerts, and contributing to dashboards. Additionally, the SRE will manage installation, patching, and upgrades of third-party applications, maintain systems aligned with N-1 patching standards, and perform ongoing maintenance. Automation is a key aspect, involving the use and maintenance of existing Ansible automation and scripts, and improving operational processes through targeted automation. Collaboration with Cloud Engineering, Development, and Security teams is essential, as is supporting the onboarding of new applications and maintaining documentation and runbooks.

Requirements

  • 4+ years of experience in Application Support, SRE, or Systems Administration
  • Experience supporting production applications in AWS
  • Hands-on experience with Deployments (Ansible / CI-CD)
  • Hands-on experience with Git workflows
  • Experience with Monitoring tools (Splunk or Google Analytics, CloudWatch)
  • Experience with Synthetic monitoring
  • Experience participating in 24/7 on-call rotations
  • Strong troubleshooting skills
  • Experience with patching and lifecycle management (N‑1)

Responsibilities

  • Provide hands-on support for enterprise applications in production
  • Troubleshoot issues across application and infrastructure layers
  • Drive incident triage and resolution
  • Ensure systems remain stable, available, and supportable
  • Participate in a 24/7 on-call rotation
  • Respond to and resolve production incidents
  • Coordinate with engineering teams during high-severity issues
  • Contribute to post-incident reviews
  • Execute deployments using Ansible and CI/CD pipelines
  • Use Git workflows for release management and rollback
  • Support validation and troubleshooting of deployment issues
  • Improve deployment reliability and processes
  • Monitor systems using Splunk or Google Analytics, CloudWatch, and Synthetic monitoring tools
  • Analyze logs, metrics, and alerts to diagnose issues
  • Contribute to dashboards and operational visibility
  • Manage installation, patching, and upgrades of third-party applications
  • Maintain systems aligned with N‑1 patching standards
  • Perform ongoing maintenance and health checks
  • Use and maintain existing Ansible automation and scripts
  • Improve operational processes through targeted automation
  • Work with Cloud Engineering, Development, and Security teams
  • Support onboarding of new applications into production
  • Maintain documentation and runbooks
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service