RESEARCH PROGRAM
Research
Our teams study the inner workings of frontier models — so that capability, safety, and societal impact can be measured rather than assumed.
Interpretability · Alignment · Evaluation · Societal Impact01
Interpretability
Reverse-engineering the computations of trained models, from single circuits to model-scale behavior.
02
Alignment
Training models whose objectives, values, and refusals remain stable as capabilities grow.
03
Evaluation
Building rigorous, pre-deployment tests for dangerous capabilities and long-horizon autonomy.
04
Societal Impact
Studying how AI systems change labor, science, and public institutions — with evidence, not anecdote.
ARCHIVE
Publications
No publications yet
Research notes will appear here when we publish them. The programmes above are the lab’s framing, not a list of papers.