About The Position

Coforge is seeking an experienced Site Reliability Engineering (SRE) Team Lead to guide our Application Support SRE function and manage a high‑performing team responsible for ensuring the performance, availability, and reliability of mission‑critical customer‑facing applications. This role combines hands‑on technical leadership with people management, process ownership, and operational excellence.

Requirements

  • 5–8+ years in SRE, DevOps, or production engineering roles.
  • 2–4+ years in a technical lead or people management capacity.
  • Strong experience supporting AWS-based applications, microservices, or API-driven environments.
  • Advanced troubleshooting skills.
  • Hands-on experience with observability stacks (either opensource or splunk)
  • Familiarity with ITIL and incident frameworks.

Nice To Haves

  • Bachelor's or Master's degree in Computer Science or related field.
  • Certifications in ITIL, AWS, Azure, or GCP.
  • Experience with Mulesoft, Postman, and API testing.
  • Proficiency with Kubernetes
  • Strong cloud-native networking knowledge.

Responsibilities

  • Manage, mentor, and coach a team of Application Support SREs; support career progression and skills development.
  • Oversee team performance, capacity planning, and staffing for a 24x7 support model.
  • Serve as the senior escalation point during major incidents and high-severity events.
  • Foster a culture of accountability, blameless postmortems, continuous learning, and operational excellence.
  • Establish team OKRs, KPIs, and reliability goals aligned with business objectives.
  • Own and mature SRE processes including incident management, problem management, change management, and service readiness.
  • Lead major incident response, coordinate cross-functional teams, ensure communication excellence, and drive root cause analysis.
  • Define and enforce SLOs, SLIs, and error budgets for supported applications.
  • Implement preventative solutions and systemic fixes that reduce incident recurrence.
  • Enhance observability practices across Splunk, OpenTelemetry, AppDynamics, Datadog, and similar tools.
  • Improve dashboards, alerting strategies, and telemetry coverage.
  • Provide insights and recommendations for reliability, scalability, and performance across AWS-hosted applications, Mulesoft APIs, and Kubernetes-based services.
  • Collaborate with development and architecture teams to integrate SRE principles early in the lifecycle.
  • Champion automation to reduce toil—CI/CD optimization, deployment improvements, self-healing mechanisms, and runbooks.
  • Analyze logs, performance issues, and code behavior to support Tier 2/Tier 3 escalations.
  • Recommend initiatives to expand Splunk automation and AI-driven insights.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service