We're looking to expand our Cloud Infrastructure team with a Senior Infrastructure Engineer, SRE to lead the reliability and operational evolution of our platform. We run hundreds of services in production, which enable us to process billions of transactions, consume multiple terabytes of data, and produce hundreds of millions of logs per day, and our reliability practice needs to evolve to match our growing scale. This includes: Building and improving the reliability and resiliency of our systems and services Establishing SLIs, SLOs, and error budgets for our most critical services and user journeys, and reviewing them regularly with the teams that own them Owning and evolving our disaster recovery strategy: recovery objectives, failover and restore paths, and regular exercises that prove they work Partnering with product engineering teams so they can own and operate their own services, with metrics that reflect real user experience Evolving our observability platform and standards across metrics, tracing, and logs: including instrumentation paved roads, alert quality, and observability cost Strengthening our incident practice: tuning paging thresholds, keeping runbooks current, and following through on postmortem action items Contributing to day-to-day Cloud Infrastructure work alongside your reliability specialty — infrastructure build-outs, platform backlog, and a shared on-call rotation (1 week out of every 6 weeks). You'll join the Cloud Infrastructure team and partner with engineering and internal support teams to drive this work. We support millions of people to improve their financial lives, and this role ensures we can continue to do so reliably and at scale.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed