Within the Tech, Data, AI, Ventures (TDAV) organization, our work is guided by a shared vision: deploying the power of technology, data, AI and ventures to accelerate sustainable competitive advantage for New York Life's businesses. We build solutions that power how we serve policy owners, agents, advisors and employees. Grounded in New York Life’s culture of long-term commitment, integrity and putting people at the center of what we do, our work reflects a purpose-driven approach to innovation. We engineer complex systems that translate into measurable business outcomes. Across technology, data, AI, cyber, product, digital experience, architecture and infrastructure, TDAV combines the scale and investment of an industry leader with the opportunity to work with leading-edge technologies and help shape how a world-class financial services company competes in the AI era — all backed by the stability and purpose of a mutual company built to last. New York Life is seeking an experienced Site Reliability Engineer (SRE) to join the Cloud Platform Engineering Team and provide both technical leadership and day-to-day management for multi cloud solution delivery (AWS, GCP, Azure). This role will partner with application teams, architects, security, operations, and business stakeholders to design, deliver, and continuously improve secure, scalable, reliable, and cost-effective solutions on AWS. The role will build new IaC artifacts and governed curated registries which will be consumed by NYL application teams. The goal is to shift from non-standard artifacts to standard platform patterns offering a full stack deployment of infrastructure which requires little to zero intervention by application development teams. Most of our applications are deployed on AWS and we leverage GCP for building agentic AI agents. Azure supports our Digital Workplace Services. The ideal candidate combines strong hands-on cloud engineering experience with the ability to guide technical direction, mentor engineers, build trusted stakeholder relationships, and drive execution across multiple workstreams. Successful reliability outcomes are likely to implement and extend on DevOps and Agile ways of working and associated automation approaches. These are underpinned by the site reliability engineer’s solid understanding of systems, production environments, operational insights, incident management, on-premises, cloud and hybrid world. The nature of the work involved means that the site reliability engineer will directly engage with customer teams but will also work on reliability initiatives that span multiple teams. The SRE collaborates closely with product owners and teams, architects, IT service management, software developers, security and network engineers, as well as other subject matter experts and roles, particularly in infrastructure and operations. Being an approachable team player and a good communicator is therefore crucial for success, and a willingness to lead initiatives is important.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed