Orientation 2026: fill out the MAIA interest form to grab merch at our events! Fill it out
Research
Organizations MAIA Works With
This is a list of some of the organizations our members have worked with.
Not all organisations listed endorse or are affiliated with MAIA.
Research by MAIA Members
Selected work coauthored by MAIA members and alumni. These projects were conducted across their respective research groups and institutions.
Highlighted papers
All research
Natural Emergent Misalignment from Reward Hacking in Production RL
Reward hacking learned in production coding environments generalized to alignment faking, malicious cooperation, and attempted safety-research sabotage.
Weight-sparse transformers have interpretable circuits
Weight sparsity produced compact, human-readable circuits; scaling improved the capability–interpretability frontier but exposed a remaining scale limit.
Distillation Robustifies Unlearning
Shows distillation can remove latent capabilities left behind by ordinary unlearning; UNDO matched retraining-level robustness with 60–80% of the compute and 0.01% labelled pretraining data, including on WMDP.
CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring
Compared CoT and action-only monitoring in adversarial coding tasks and tested when harmful side goals remain visible to a trusted monitor.
On the creation of narrow AI: hierarchy and nonlocality of neural network skills
Finds that broad curricula may be required to learn narrow skills and that skills are not perfectly localisable, while pruning-based transfer can still outperform distillation.
Scaling Laws For Scalable Oversight
Built and validated a quantitative oversight-scaling model across Nim, Mafia, Debate, Backdoor Code, and Wargames.
Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
A weaker monitor detected reward hacking from a stronger model's chain of thought, while optimizing against the monitor encouraged hidden intent rather than eliminating most misbehavior.
Harmonic Loss Trains Interpretable AI Models
Studies an alternative training loss designed to improve model interpretability.
International AI Safety Report
Synthesizes the evidence on general-purpose AI capabilities, systemic risks, evaluations, and safeguards for an international policy audience.
Diverse Preference Learning for Capabilities and Alignment
Studies preference learning methods that preserve diversity rather than optimizing for a single preferred response.
Alignment faking in large language models
Claude 3 Opus selectively complied during training; harmful-query compliance reached 14% in the training-signalled condition and explicit alignment-faking reasoning rose after RL.
Efficient Dictionary Learning with Switch Sparse Autoencoders
Switch SAEs routed activations across expert autoencoders and delivered a substantial reconstruction-versus-sparsity Pareto improvement at fixed compute.
Unlearning-based Neural Interpretations
Introduces an adaptive unlearning-based attribution baseline that removes salient features, smooths local decision boundaries, and produces more faithful and robust interpretations than static baselines.
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Turned 75 Kaggle competitions into an agent benchmark; the best evaluated setup reached bronze-medal performance on 16.9% of competitions.
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
Introduced a broad benchmark for whether LLMs know what they are and the circumstances in which they operate, including evaluation-versus-deployment distinctions.
Not All Language Model Features Are One-Dimensionally Linear
Found irreducible circular features for concepts such as weekdays and months and causally linked them to modular computations in multiple LLMs.
An Assessment of Model-on-Model Deception
Evaluates whether one language model can detect deception by another.
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
Introduced causally implicated, human-interpretable feature circuits; used them to improve classifier generalization and automate circuit discovery at scale.
Building an early warning system for LLM-aided biological threat creation
Evaluates how language-model assistance affects the ability to carry out biological-threat tasks.
Black-Box Access is Insufficient for Rigorous AI Audits
Argues from concrete audit failure modes that query-only access cannot support rigorous external audits and specifies stronger access requirements.
A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Found that DPO suppressed toxic outputs by bypassing rather than removing capabilities, then used the mechanism to reverse the alignment behavior.
Model Manipulation Attacks Enable More Rigorous Evaluations of LLM Capabilities
Tests whether modifying a model can expose capabilities missed by ordinary evaluations.
Forbidden Facts: An Investigation of Competing Objectives in Llama-2
Tests how language models respond when instructions to withhold information conflict with other objectives.
Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
Strong models exceeded weak supervisors, and an auxiliary confidence loss recovered nearly 80% of the GPT-2-to-GPT-4 performance gap on NLP tasks.
Open Problems and Fundamental Limitations of RLHF
A field-level taxonomy of RLHF failure modes, complementary methods, and auditing standards.
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
Evaluated GPT-3.5 and GPT-4 across toxicity, bias, robustness, privacy, ethics, and fairness, uncovering jailbreak and data-leakage vulnerabilities.
Explore, Establish, Exploit: Red Teaming Language Models from Scratch
Studies how red-team attacks can be developed against language models without an existing attack dataset.
The Quantization Model of Neural Scaling
Models neural scaling as the acquisition of discrete skills and tests the resulting predictions.
Red Teaming Deep Neural Networks with Feature Synthesis Tools
Introduces a 12-trojan benchmark for whether interpretability tools help humans find unknown model bugs; evaluates 16 attribution tools and seven feature-synthesis methods.
Adversarial Policies Beat Superhuman Go AIs
A learned adversary beat superhuman KataGo more than 97% of the time; the exploit transferred and could be reproduced by human experts.
Scaling Laws for Reward Model Overoptimization
Measured Goodhart-style overoptimization in RLHF proxies and found smooth scaling relationships across model, data, and optimization choices.
Locating and Editing Factual Associations in GPT
Localises factual recall to mid-layer feed-forward computations and introduces ROME, which edits specific facts while preserving specificity and generalisation.
Robust Feature-Level Adversaries are Interpretability Tools
Built targeted, universal, disguised, physically realizable, and black-box feature-level attacks and used them to predict natural copy-paste failures.










