From a reliability standpoint, this role involves evaluating the scalability, resiliency, performance, and security properties and techniques used in production environments. It supports the uptime of production services through an On-Call rotation, which includes monitoring and alerting to meet internal Service Level Objectives (SLOs) and customer-facing Service Level Agreements (SLAs). Ensuring reliable incident processes is achieved by conducting Disaster Recovery drills. The role also focuses on improving reliability through incident management by investigating incidents, implementing remediation strategies, and learning from past incidents to make improvements. It involves determining the reliability and security requirements of components and systems to meet the reliability objectives of the company, customers, and any relevant governmental agencies. Additionally, the role aims to reduce operational expenses through automation, by identifying and mitigating failure points, and automating repetitive and resource-intensive tasks. It also involves developing new acceleration techniques and analytical tools to ensure the early identification of potential issues with new products, packaging, processes, and overall product reliability. Ensures a reliable and scalable network, manages network / cloud infrastructure and storage systems supporting business operations, and responds to planned maintenance, real-time outages, and issues. Plans, designs and implements local and wide-area network solutions between multiple platforms and protocols.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior