This role involves developing software solutions to enable the reliability and operability of large-scale distributed systems. The engineer will build a deep understanding of system behavior, scaling, interaction, and failure points to identify risks and opportunities for remediation. Key responsibilities include implementing monitoring and reporting for production environments, building tools and automation to eliminate toil and reduce operational overhead, and creating frameworks, processes, and best practices for the engineering team. The role also involves defining meaningful Service Level Indicators (SLIs), automating critical engineering processes to minimize risk and maximize innovation speed, and managing capacity and performance to scale infrastructure on both public and private clouds. A deep dive into understanding application components to promote product scalability, stability, and performance, along with continuous delivery, performance fine-tuning, and troubleshooting, are also key aspects of this position.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed