Delivers services at high scale and high availability with resilience by using automation and Infrastructure Code. Builds reliability into the ecosystem by applying best practices in resiliency engineering and observability by developing resiliency tools and capabilities for observability and chaos testing of pipelines. Combines systems and software engineering techniques with site reliability engineering practices to create reliable user experiences to support workplace investing, healthcare and defined benefits organizations. Assists teams scale through production insights, operational automation, developer guidance, real-time metrics, and automation. Provides training, support, and alignment to ensures Site Reliability Engineers have the skills, tools, and opportunities to accomplish engineering reliable systems. Provides Product and Platform teams with engineering expertise that enables them to clarify and achieve their system reliability goals. Partners with our key stakeholders in defining and adopting policies, processes and practices that lead to reliable Information Technology systems and measures compliance with those policies.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Principal