This role is a Technical Lead position focused on Cloud & Infrastructure Engineering, with a primary area of responsibility in Site Reliability Engineering (SRE). The position involves ensuring the high availability, performance, and resilience of production systems by implementing SRE best practices such as error budgets, SLIs/SLOs, capacity planning, chaos testing, and runbook creation. A key aspect of the role is driving automation to reduce manual operational tasks and improve Mean Time To Recovery (MTTR). The Technical Lead will also conduct post-incident reviews (PIRs) and implement long-term corrective actions. In Middleware & Application Platform Management, responsibilities include managing deployments, rollbacks, and environment synchronization across different stages (Dev, QA, UAT, Production). This involves installing, configuring, upgrading, and maintaining WebLogic, Tomcat, Apache, and Nginx servers, along with performing JVM tuning, thread pool optimization, connection pool management, and general performance tuning. Troubleshooting middleware issues such as memory leaks, thread contention, SSL, certificates, and clustering is also a core function. The role also encompasses Monitoring, Logging & Observability, requiring configuration of dashboards, alerts, and performance insights using tools like New Relic and Splunk. Developing log-based monitoring strategies, anomaly detection, and implementing proactive monitoring to reduce downtime are essential. CDN & Edge Platform Management, specifically with Akamai, involves configuring caching rules, WAF policies, edge redirects, and performance optimizations. Troubleshooting CDN-related latency, caching, and routing issues, and collaborating with Akamai support for advanced issues are part of the duties. In Incident, Problem & Change Management, the Technical Lead will lead major incident bridges, coordinate cross-functional teams, and provide timely updates. Managing problem tickets, conducting root cause analysis, and developing preventive action plans are crucial. Ensuring compliance with ITIL processes for change, release, and incident management is also required. Leadership & Stakeholder Management involves leading and mentoring a team of SRE/DevOps engineers, providing technical guidance, training, and performance feedback. The role requires acting as a customer-facing technical SME for escalations and production issues, and collaborating with product, QA, development, and business teams to ensure smooth delivery.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed