The Safeguards ML Infra team designs, builds, and operates the production infrastructure that powers Claude's safety systems. They own the critical backend services that ensure safety on the token generation path and the operational work of getting those systems safely into production. This includes standing up safeguards for every new model launch and deploying new safety classifiers as they ship. Every frontier model release runs through this team, configuring, verifying, and rolling out safeguards across all platforms Claude runs on (1P, AWS Bedrock, GCP Vertex, etc.). The team also leads incident response when issues arise. This role is central to that operational work. The engineer will ensure safeguards are properly configured and deployed for model launches and own the off-cycle deployment of new safety classifiers. This involves canarying changes, verifying that the right safeguards are provably live on the right models, and holding rollback authority when issues occur. The goal is to evolve manual verifications into a self-running system, turning launch runbooks into tooling, hand-built checks into continuous validation, and one-off deploys into a repeatable pipeline. The ideal candidate has deep experience in production change management at scale, has owned deploy pipelines, config management systems, rollout safety, or launch readiness for systems under real production pressure. Familiarity with ML research or transformer architectures is not required, as this will be learned on the job. The priority is production judgment: a track record of shipping changes to critical systems safely and automating oneself out of previous work.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
Associate degree