Research Scientist, Applied White-Box Methods

FAR.AI•Berkeley, CA
•$150,000 - $250,000•Onsite

About The Position

The Applied White-Box Methods team develops, evaluates, and demonstrates methods that leverage model internals to improve the safety of AI systems. We work on diverse AI safety applications of white-box methods, from white-box control to evaluation awareness to shaping training dynamics to improve alignment. Black-box methods, such as chain-of-thought monitoring, work well for now, but we are quickly entering a world where black-box interventions and monitoring are insufficient. Interpretability research is still often early-stage, curiosity-driven work without realistic evaluations on applications that matter. The team bridges the gap between exploration and deployment by stress-testing white-box methods on real-world (e.g. long-context agentic coding) tasks and using this feedback loop to enable the deployment of better white-box methods at frontier scale. Our scope includes any method that uses model internals to understand, predict, or intervene on model behavior, not only what is conventionally called interpretability. Methods of current interest include natural language autoencoders and other activation explainers, activation oracles, steering, patching, and other activation-level interventions, influence functions and data attribution, and singular learning theory. We evaluate our methods against real baselines – strong black-box methods and activation probes – to be able to make an honest case that the methods are worth implementing, or conclude that simpler methods work better for now. AI Research Automation. AI will soon automate most of the hill-climbing in the research process. We anticipate this by focusing our effort and judgement on defining realistic evaluations with Goodhart-resistant metrics. With the evaluation framework properly built, we can pour vast amounts of AI labor into method development and iteration without overfitting. We believe this is the way to scale white-box research into the age of RSI. Realistic, large-scale models. Unrealistic models lead to unconfident conclusions about which methods do and do not work. We leverage FAR's shared compute and infrastructure to work on organisms that come out of pipelines a frontier lab could plausibly have run: for example, reward seekers trained by RL in broken environments, with other contaminated data mixed in to induce other misalignments. Realistic, large-scale evaluations. We will primarily study long-context agentic coding as this is where most of the risk currently lies. In addition to the classic AI control sabotage settings, we will also study reward hacking, sandbagging, and research tampering, which are some of the key failure modes that matter during RSI. Practical monitors and interventions. Deployment of new methods has real costs for AI developers. Part of our focus on real-world applications is ensuring that methods are simple and efficient enough to deploy at frontier scale.

Requirements

  • Researchers with some experience in more fundamental interpretability who want to evaluate and refine those methods in realistic settings, including agentic coding, long contexts, and realistic threat models.
  • Researchers from evaluations, AI control, red-teaming, or reinforcement learning who have begun working with model internals and want to deepen that work.
  • Prior hands-on experience with white-box methods and a demonstrated interest in their practical application are preferred.
  • Be prepared to explain how your previous research experience (e.g. in applied ML) could be leveraged for our work, and how you are engaging with the field of technical AI safety research today.
  • Hands-on experience applying at least one white-box method to a real model (activation explainers, SAEs, steering, attribution, influence functions, probes, or similar), and an informed view of its limitations.
  • A track record in AI safety: a paper, a fellowship project, or substantive public writing.
  • Experience with evaluations, AI control, red-teaming, reinforcement learning, or post-training of LLMs.
  • The ability to communicate novel methods and results clearly to technical and non-technical audiences.
  • A PhD or several years of research experience in computer science, machine learning, physics, statistics, or a related field.
  • Previous experience in applied ML for other fields: e.g. biology, chemistry, materials science, robotics, etc.

Nice To Haves

  • Alumni of programs such as MATS, Astra, Anthropic Fellows, SPAR, Pivotal, LASR or similar programs are especially encouraged to apply.
  • Experience running large training runs is a plus.

Responsibilities

  • Take ownership of and accelerate the team's research agenda.
  • Publish findings broadly.
  • Engage with the AI alignment community.
  • Propose new directions within the team's agenda.
  • Attend relevant conferences and other community events.
  • Leverage FAR.AI’s existing comprehensive infrastructure for events convening and government relations.
  • Work with national AI safety institutes, frontier model developers, and top academics.

Benefits

  • Health Insurance - 94% of Insurance premium paid by Organization commencing within 1 month after your start date
  • Retirement - 401(k) plan with up to 2% match
  • PTO - 25 days Paid Time Off per year, accrued weekly and up to 10 days of paid sick leave per year
  • Paid Leave - Paid Bereavement, Family, Medical and Pregnancy Disability Leave
  • WFH Stipend & Equipment - Work computer and stipend provided for eligible employees
  • Catered Meals (Berkeley Office Only) - Catered lunches and dinners on workdays at our office

Stand Out From the Crowd

Upload your resume and get instant feedback on how well it matches this job.

Upload and Match Resume

What This Job Offers

Job Type

Full-time

Career Level

Entry Level

Education Level

Ph.D. or professional degree

© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service