This role serves as the highest level technical escalation point (L3 SME) for Linux/Unix infrastructure. The position involves managing and supporting large-scale Red Hat Enterprise Linux (RHEL), Oracle Linux, CentOS, SUSE, and Unix environments. Key responsibilities include advanced troubleshooting of OS, kernel, filesystem, storage, networking, and performance issues, as well as leading OS patching, upgrades, vulnerability remediation, and lifecycle management. The role ensures system availability, stability, scalability, and operational compliance, while also leading the resolution of critical incidents and major outages. Additionally, the position involves performing root cause analysis (RCA), implementing preventive measures, reviewing and approving changes related to compute infrastructure, and driving continuous service improvement initiatives. Capacity planning, resource optimization, analyzing system utilization, bottlenecks, and trends are also crucial. The role will implement proactive monitoring and self-healing mechanisms, and drive toil reduction through automation and operational innovation. Developing and maintaining shell scripting and automation solutions, automating provisioning, patching, compliance checks, and operational tasks are key. Contribution to Infrastructure as Code practices using Terraform and supporting CI/CD integration for infrastructure deployment activities are also expected.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed