Orientation 2026: fill out the MAIA interest form to grab merch at our events! Fill it out

Research

Organizations MAIA Works With

This is a list of some of the organizations our members have worked with.
Not all organisations listed endorse or are affiliated with MAIA.

Research by MAIA Members

Selected work coauthored by MAIA members and alumni. These projects were conducted across their respective research groups and institutions.

Weight-sparse transformers have interpretable circuits

Weight sparsity produced compact, human-readable circuits; scaling improved the capability–interpretability frontier but exposed a remaining scale limit.

MAIA coauthors: Leo Gao, Achyuta Rajaram

Distillation Robustifies Unlearning

Shows distillation can remove latent capabilities left behind by ordinary unlearning; UNDO matched retraining-level robustness with 60–80% of the compute and 0.01% labelled pretraining data, including on WMDP.

MAIA coauthors: Leni Shor

Scaling Laws For Scalable Oversight

Built and validated a quantitative oversight-scaling model across Nim, Mafia, Debate, Backdoor Code, and Wargames.

MAIA coauthors: Josh Engels, David D. Baek

International AI Safety Report

Synthesizes the evidence on general-purpose AI capabilities, systemic risks, evaluations, and safeguards for an international policy audience.

MAIA coauthors: Tamay Besiroglu, Stephen Casper

Alignment faking in large language models

Claude 3 Opus selectively complied during training; harmful-query compliance reached 14% in the training-signalled condition and explicit alignment-faking reasoning rose after RL.

MAIA coauthors: Benjamin Wright (MAIA Alum)

Unlearning-based Neural Interpretations

Introduces an adaptive unlearning-based attribution baseline that removes salient features, smooths local decision boundaries, and produces more faithful and robust interpretations than static baselines.

MAIA coauthors: Ching Lam Choi

Black-Box Access is Insufficient for Rigorous AI Audits

Argues from concrete audit failure modes that query-only access cannot support rigorous external audits and specifies stronger access requirements.

MAIA coauthors: Stephen Casper, Marvin von Hagen, Misha Gerovitch, Wendy Sun

Adversarial Policies Beat Superhuman Go AIs

A learned adversary beat superhuman KataGo more than 97% of the time; the exploit transferred and could be reproduced by human experts.

MAIA coauthors: Tony Wang

Scaling Laws for Reward Model Overoptimization

Measured Goodhart-style overoptimization in RLHF proxies and found smooth scaling relationships across model, data, and optimization choices.

MAIA coauthors: Leo Gao

Locating and Editing Factual Associations in GPT

Localises factual recall to mid-layer feed-forward computations and introduces ROME, which edits specific facts while preserving specificity and generalisation.

MAIA coauthors: Kevin Meng