This role involves working alongside development and operations teams to ensure speedy and reliable software deployments, monitor systems, and enhance platform reliability. The engineer will be responsible for discovering, documenting, and fixing system bugs. Key responsibilities include developing features using AI coding tools and script repositories to automate, scale, test, and secure cloud infrastructure and pipelines. The position requires enhancing performance monitoring through tools like Splunk, identifying and optimizing performance bottlenecks, and contributing to the SRE journey by suggesting improvements in engineering build, maintenance, automation, and reliability using SRE/DevOps tools and Infrastructure-as-Code. The role also includes developing and coding high-quality pipeline automation workflows, creating and executing test strategies for failure scenarios, and building automated systems for continuous performance, stress, and load testing. Collaboration with SREs, developers, and operations teams to define reliability goals and testing strategies is essential. The engineer will ensure new services and features are thoroughly tested for performance, reliability, and failure recovery before production deployment, and validate monitoring, logging, and alerting mechanisms. Measuring and tracking Service Level Indicators (SLIs) and Service Level Objectives (SLOs) through automated testing frameworks is also a key part of the role. The engineer is expected to resolve most conflicts between timeline, budget, and scope independently, while escalating complex issues to senior management.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level