This role involves working closely with development and operations teams to ensure speedy and reliable software deployments, monitor systems, and enhance platform reliability. The engineer will be responsible for identifying and fixing system bugs, developing features using AI coding tools and script repositories for automation, scaling, testing, and securing cloud infrastructure and pipelines. A key aspect of the role is enhancing performance monitoring through tools like Splunk, identifying and optimizing performance bottlenecks, and contributing to the SRE journey by improving engineering build, maintenance, automation, and reliability using SRE/DevOps tools and Infrastructure-as-Code. The position requires developing and coding high-quality pipeline automation workflows, creating and executing test strategies for failure scenarios, and building automated systems for continuous performance, stress, and load testing. Collaboration with SREs, developers, and operations teams to define reliability goals and testing strategies is essential. The role also includes ensuring new services and features are thoroughly tested before production, validating monitoring, logging, and alerting mechanisms, and ensuring accurate measurement and tracking of SLIs and SLOs. The engineer will independently resolve most conflicts between timeline, budget, and scope, escalating complex issues to senior management.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level