Intern: AI Red Teaming (Fall 2026)

Realm LabsSunnyvale, CA

About The Position

You will try to break the systems we build and the models we protect, and turn what you find into something the team can act on: a reproducible attack, an evaluation that catches it, and a written account of why it works. Expect some mix of: eliciting unsafe behaviour from aligned LLMs and multi-modal models; prompt injection and tool-use abuse against agentic systems; automating attack generation and evaluation rather than hand-crafting one-off prompts; and measuring whether guardrails hold under pressure. Where RealmLabs' interpretability work gives you access to a model's internals, use it. We aim for a paper or public technical report out of every internship, plus attacks that stay in our evaluation suite after you leave.

Requirements

  • Hands-on experience attacking or stress-testing models (jailbreaks, prompt injection, adversarial examples, data poisoning, model extraction, or evaluating safety and moderation systems)
  • Able to read a paper and implement its attack
  • Machine learning tools: pytorch, huggingface, transformers, datasets
  • Applied deep learning and LLM experience
  • Training and evaluating deep models
  • Development environments and tools: unix, git, basic clouds usage on AWS and/or GCP, jupyter
  • Programming: python
  • Must be authorized to work in the USA or must be able to obtain CPT (Curricular Practical Training) approval from host university.

Nice To Haves

  • Offensive security background outside ML: CTFs, vulnerability research, penetration testing
  • Familiarity with agentic systems and their attack surface — tool calls, retrieval, memory, multi-agent orchestration
  • Finetuning LLMs, multi-modal LLMs
  • Familiarity with ML[NLP,LLM,Vision] interpretability methods, sparse autoencoders, linear probes
  • Programming languages well-roundedness
  • Experience in statically-typed and functional languages

Responsibilities

  • Eliciting unsafe behaviour from aligned LLMs and multi-modal models
  • Prompt injection and tool-use abuse against agentic systems
  • Automating attack generation and evaluation
  • Measuring whether guardrails hold under pressure
  • Using model internals to locate failure modes

Benefits

  • Market aligned compensation for interns in the bay area
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service