Skip to content
RESEARCH PROGRAM

Research

Our teams study the inner workings of frontier models — so that capability, safety, and societal impact can be measured rather than assumed.

Interpretability · Alignment · Evaluation · Societal Impact
01

Interpretability

Reverse-engineering the computations of trained models, from single circuits to model-scale behavior.

02

Alignment

Training models whose objectives, values, and refusals remain stable as capabilities grow.

03

Evaluation

Building rigorous, pre-deployment tests for dangerous capabilities and long-horizon autonomy.

04

Societal Impact

Studying how AI systems change labor, science, and public institutions — with evidence, not anecdote.

ARCHIVE

Publications

No publications yet

Research notes will appear here when we publish them. The programmes above are the lab’s framing, not a list of papers.