About The Position

This role is a Technical Lead position focused on Cloud & Infrastructure Engineering, with a primary area of responsibility in Site Reliability Engineering (SRE). The position involves ensuring the high availability, performance, and resilience of production systems by implementing SRE best practices such as error budgets, SLIs/SLOs, capacity planning, chaos testing, and runbook creation. A key aspect of the role is driving automation to reduce manual operational tasks and improve Mean Time To Recovery (MTTR). The Technical Lead will also conduct post-incident reviews (PIRs) and implement long-term corrective actions. In Middleware & Application Platform Management, responsibilities include managing deployments, rollbacks, and environment synchronization across different stages (Dev, QA, UAT, Production). This involves installing, configuring, upgrading, and maintaining WebLogic, Tomcat, Apache, and Nginx servers, along with performing JVM tuning, thread pool optimization, connection pool management, and general performance tuning. Troubleshooting middleware issues such as memory leaks, thread contention, SSL, certificates, and clustering is also a core function. The role also encompasses Monitoring, Logging & Observability, requiring configuration of dashboards, alerts, and performance insights using tools like New Relic and Splunk. Developing log-based monitoring strategies, anomaly detection, and implementing proactive monitoring to reduce downtime are essential. CDN & Edge Platform Management, specifically with Akamai, involves configuring caching rules, WAF policies, edge redirects, and performance optimizations. Troubleshooting CDN-related latency, caching, and routing issues, and collaborating with Akamai support for advanced issues are part of the duties. In Incident, Problem & Change Management, the Technical Lead will lead major incident bridges, coordinate cross-functional teams, and provide timely updates. Managing problem tickets, conducting root cause analysis, and developing preventive action plans are crucial. Ensuring compliance with ITIL processes for change, release, and incident management is also required. Leadership & Stakeholder Management involves leading and mentoring a team of SRE/DevOps engineers, providing technical guidance, training, and performance feedback. The role requires acting as a customer-facing technical SME for escalations and production issues, and collaborating with product, QA, development, and business teams to ensure smooth delivery.

Requirements

  • Ownership & Accountability: Takes responsibility for production stability and issue resolution.
  • Leadership: Guides team members, manages workload, and drives operational excellence.
  • Communication: Clear, structured communication with customers and internal teams.
  • Problem Solving: Strong analytical skills and ability to troubleshoot complex issues.
  • Collaboration: Works effectively across engineering, QA, product, and business teams.
  • Calm Under Pressure: Handles critical incidents with composure and clarity.

Responsibilities

  • Ensure high availability, performance, and resilience of production systems.
  • Implement SRE best practices: error budgets, SLIs/SLOs, capacity planning, chaos testing, runbook creation.
  • Drive automation to reduce manual operational tasks and improve MTTR.
  • Conduct post‑incident reviews (PIRs) and implement long‑term corrective actions.
  • Manage deployments, rollbacks, and environment synchronization across Dev, QA, UAT, and Production.
  • Install, configure, upgrade, and maintain WebLogic, Tomcat, Apache, and Nginx servers.
  • Perform JVM tuning, thread pool optimization, connection pool management, and performance tuning.
  • Troubleshoot middleware issues related to memory leaks, thread contention, SSL, certificates, and clustering.
  • Configure dashboards, alerts, and performance insights using New Relic and Splunk.
  • Develop log‑based monitoring strategies and anomaly detection.
  • Implement proactive monitoring to reduce downtime and improve reliability.
  • Configure Akamai caching rules, WAF policies, edge redirects, and performance optimizations.
  • Troubleshoot CDN‑related latency, caching, and routing issues.
  • Collaborate with Akamai support for advanced troubleshooting.
  • Lead major incident bridges, coordinate cross functional teams, and provide timely updates.
  • Manage problem tickets, root cause analysis, and preventive action plans.
  • Ensure compliance with ITIL processes for change, release, and incident management.
  • Lead and mentor a team of SRE/DevOps engineers.
  • Provide technical guidance, training, and performance feedback.
  • Act as a customer facing technical SME for escalations and production issues.
  • Collaborate with product, QA, development, and business teams to ensure smooth delivery.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service