The Applied White-Box Methods team develops, evaluates, and demonstrates methods that leverage model internals to improve the safety of AI systems. We work on diverse AI safety applications of white-box methods, from white-box control to evaluation awareness to shaping training dynamics to improve alignment. Black-box methods, such as chain-of-thought monitoring, work well for now, but we are quickly entering a world where black-box interventions and monitoring are insufficient. Interpretability research is still often early-stage, curiosity-driven work without realistic evaluations on applications that matter. The team bridges the gap between exploration and deployment by stress-testing white-box methods on real-world (e.g. long-context agentic coding) tasks and using this feedback loop to enable the deployment of better white-box methods at frontier scale. Our scope includes any method that uses model internals to understand, predict, or intervene on model behavior, not only what is conventionally called interpretability. Methods of current interest include natural language autoencoders and other activation explainers, activation oracles, steering, patching, and other activation-level interventions, influence functions and data attribution, and singular learning theory. We evaluate our methods against real baselines – strong black-box methods and activation probes – to be able to make an honest case that the methods are worth implementing, or conclude that simpler methods work better for now. AI Research Automation. AI will soon automate most of the hill-climbing in the research process. We anticipate this by focusing our effort and judgement on defining realistic evaluations with Goodhart-resistant metrics. With the evaluation framework properly built, we can pour vast amounts of AI labor into method development and iteration without overfitting. We believe this is the way to scale white-box research into the age of RSI. Realistic, large-scale models. Unrealistic models lead to unconfident conclusions about which methods do and do not work. We leverage FAR's shared compute and infrastructure to work on organisms that come out of pipelines a frontier lab could plausibly have run: for example, reward seekers trained by RL in broken environments, with other contaminated data mixed in to induce other misalignments. Realistic, large-scale evaluations. We will primarily study long-context agentic coding as this is where most of the risk currently lies. In addition to the classic AI control sabotage settings, we will also study reward hacking, sandbagging, and research tampering, which are some of the key failure modes that matter during RSI. Practical monitors and interventions. Deployment of new methods has real costs for AI developers. Part of our focus on real-world applications is ensuring that methods are simple and efficient enough to deploy at frontier scale.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Entry Level
Education Level
Ph.D. or professional degree