Infrastructure Engineers build the foundation for Ivo’s entire platform. Customers are cagey about their contracts, so each customer gets their isolated environment with containers, database, VPC, etc. Things break. Regions go down. Cloud and LLM providers have “incidents.” Customers still expect us to hit our SLAs. We’re looking for an Senior or Staff Site level Reliability Engineer as part of Infrastructure team to: Own uptime, reliability, and performance end-to-end Define and enforce SLI, SLO and SLA targets (and make sure we don’t get paged in the wee hours) Design failover + disaster recovery that actually works in real scenarios Turn data residency requirements into real systems (geo-fencing, regional isolation, etc.) Implement security controls that pass audits and don’t slow the product to a crawl Build observability that answers: what, why and how often it broke ? Lead incident response + write postmortems that make people actually read. We need someone who: Minimum 7 years of experience Thinks in failure modes Designs systems that keep working anyway Can translate “this clause in a contract” into actual infrastructure constraints This isn’t a “keep the lights on” role. You’ll be building the system that keeps the company running. In addition to helping us run a solid, high-performance distributed system, we’d love someone who’s as excited about LLMs as we are. You’d be deeply embedded into the engineering team and highly encouraged to push the frontiers.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed