Events, programs, and opportunities from MAIA. Join mailing list
MIT Faculty and Labs
Faculty and Labs
A good way to get into AI safety is to do a UROP at one of the MIT labs doing AI safety work. This can be a good alternative to fellowships because you can more easily UROP while taking classes. Below are some of the labs whose research is most relevant!
Algorithmic Alignment Group (Dylan Hadfield-Menell)
CSAIL · Value alignment, human-AI interaction, AI policy
The Algorithmic Alignment Group studies the agent alignment problem: how to identify AI behavior that is consistent with the goals of individual users, human-AI teams, multi-agent systems, and society at large. The group combines conceptual work, algorithm design, and policy, with research spanning preference and value learning, incentives in recommender systems, the limitations of RLHF, and red-teaming and robustness for language models. Much of the foundational formal work on corrigibility and reward misspecification in AI safety came out of this line of research.
Notable work
- The Off-Switch Game — Hadfield-Menell, Dragan, Abbeel & Russell, IJCAI 2017
- Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback — Casper, Davies et al., TMLR 2023
- Prompt Injection as Role Confusion — Ye, Cui & Hadfield-Menell, 2026
Tegmark AI Safety Group (Max Tegmark)
Physics, IAIFI · Formal verification, mechanistic interpretability
Max Tegmark’s group works at the intersection of physics and machine learning (using AI for physics and physics for AI), with a focus on AI safety and mechanistic interpretability. Their interpretability work reverse-engineers how neural networks represent knowledge, including the discovery that language models learn literal linear "world models" of space and time, the geometric structure of the concept space that sparse autoencoders recover, and effective theories of phenomena like grokking. The group now works largely on formal verification and "guaranteed safe AI": using formal methods and mathematical proof to obtain hard guarantees about what an AI system can and cannot do, rather than relying on empirical testing alone. Tegmark is additionally the co-founder and president of the Future of Life Institute.
Notable work
- Language Models Represent Space and Time — Gurnee & Tegmark, ICLR 2024
- The Geometry of Concepts: Sparse Autoencoder Feature Structure — Li, Michaud, Baek, Engels, Sun & Tegmark, 2024
- Towards Understanding Grokking: An Effective Theory of Representation Learning — Liu, Kitouni, Nolte, Michaud, Tegmark & Williams, NeurIPS 2022
Isola Lab (Phillip Isola)
EECS, CSAIL · Representation learning, vision, emergent intelligence
Phillip Isola’s group studies the fundamental principles of intelligence, with a focus on how neural networks represent the world and how general-purpose intelligence emerges from embodied interaction rather than imitation. Their work spans computer vision, self-supervised and contrastive representation learning, and generative models. One current thrust is "representational universals": the Platonic Representation Hypothesis argues that models trained on different datasets, objectives, and even different modalities are all converging toward a shared statistical model of reality, which has interesting implications for interpretability and for how we should expect future systems to represent the world.
Notable work
- The Platonic Representation Hypothesis — Huh, Cheung, Wang & Isola, ICML 2024
Language & Intelligence Group (LINGO) (Jacob Andreas)
EECS, CSAIL · NLP, interpretability, in-context learning
Jacob Andreas’s group studies language as a communicative and computational tool: how machines can learn from and communicate with humans in natural language, and how language can be used to understand what neural networks learn. Much of their work bears directly on interpretability, including automatically describing the features neural networks learn in natural language, probing the implicit world models inside language models, and understanding the mechanisms behind in-context learning.
Notable work
- What learning algorithm is in-context learning? Investigations with linear models — Akyürek, Schuurmans, Andreas, Ma & Zhou, ICLR 2023
- Natural Language Descriptions of Deep Visual Features — Hernandez, Schwettmann, Bau, Bagashvili, Torralba & Andreas, ICLR 2022
- Implicit Representations of Meaning in Neural Language Models — Li, Nye & Andreas, ACL 2021
(Neil Thompson)
CSAIL, Sloan · Economics of computing, AI progress and risk
MIT FutureTech is an interdisciplinary group studying the foundations of progress in computing: how trends in hardware, algorithms, and AI drive (and constrain) economic growth and productivity. Their quantitative work measures the compute and algorithmic efficiency behind recent AI capability gains, and argues that continued progress on the current paradigm requires enormous compute growth. The lab also runs the MIT AI Risk Initiative, whose AI Risk Repository is a comprehensive taxonomy and living database of AI risks now used by the UN, the EU AI Office, and the UK AI Security Institute.
Notable work
- Algorithmic Progress in Language Models — Ho, Besiroglu, Erdil et al. (with Thompson), NeurIPS 2024
- The AI Risk Repository: A Comprehensive Meta-Review, Database, and Taxonomy of Risks from AI — Slattery et al., 2024 · browse the database at airisk.mit.edu
- The Computational Limits of Deep Learning — Thompson, Greenewald, Lee & Manso, 2020
Sculpting Evolution (Kevin Esvelt)
Media Lab · Biosecurity, pandemic prevention, AI-bio risk
The Sculpting Evolution group works to advance biotechnology safely, applying molecules, models, and cryptography to defend against pandemics and prevent the catastrophic misuse of biotechnology. Esvelt was the first to recognize that CRISPR-based gene drives could alter wild populations, and published the risks alongside a call for safeguards before anyone built one. The group’s biosecurity portfolio includes the "Delay, Detect, Defend" pandemic-defense framework, SecureDNA (free cryptographic screening of DNA synthesis orders worldwide), and the Nucleic Acid Observatory for metagenomic early warning. It also produced some of the earliest empirical work at the AI-biosecurity intersection, showing that chatbots can walk non-experts through acquiring pandemic-capable pathogens.
Notable work
- Can Large Language Models Democratize Access to Dual-Use Biotechnology? — Soice, Rocha, Cordova, Specter & Esvelt, 2023
- Will Releasing the Weights of Future Large Language Models Grant Widespread Access to Pandemic Agents? — Gopal et al., 2023
- Delay, Detect, Defend: Preparing for a Future in which Thousands Can Release New Pandemics — Esvelt, GCSP Geneva Paper 29/22, 2022