You will try to break the systems we build and the models we protect, and turn what you find into something the team can act on: a reproducible attack, an evaluation that catches it, and a written account of why it works. Expect some mix of: eliciting unsafe behaviour from aligned LLMs and multi-modal models; prompt injection and tool-use abuse against agentic systems; automating attack generation and evaluation rather than hand-crafting one-off prompts; and measuring whether guardrails hold under pressure. Where RealmLabs' interpretability work gives you access to a model's internals, use it. We aim for a paper or public technical report out of every internship, plus attacks that stay in our evaluation suite after you leave.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Career Level
Intern
Education Level
No Education Listed