AI NEWS PAGE DOSSIER
Chris Olah
Chris Olah's file: neural-network visualization, mechanistic interpretability, Anthropic, and the attempt to understand what giant models are doing inside.
Olah spent years inventing ways to see inside neural networks, moved from Google Brain and Distill into OpenAI's circuits work, then led transformer-circuits and large-model interpretability research at Anthropic.
HOW THE HELL DID CHRIS OLAH GET HERE?
Sep 8, 2016
HE STARTS DRAWING THE MACHINE'S INTERNAL LOGIC
At Google Brain, Olah published a visual exploration of attention and augmented recurrent neural networks, part of a broader effort to make neural systems understandable by inspection.
RECEIPTS: Distillfirst-party
Nov 7, 2017
HE TURNS NEURONS INTO PICTURES PEOPLE CAN READ
Olah and collaborators published 'Feature Visualization,' a widely used framework for showing what image-network units respond to across layers.
RECEIPTS: Distillfirst-party
Apr 14, 2020
OPENAI PUTS THE CIRCUITS UNDER A MICROSCOPE
Olah contributed to OpenAI Microscope, a collection of visualizations meant to let researchers inspect the internal features and circuits of vision models.
RECEIPTS: OpenAIfirst-party
2021
INTERPRETABILITY MOVES INTO TRANSFORMERS
Olah and collaborators published a mathematical framework for transformer circuits, extending mechanistic-interpretability methods from vision systems into attention-based language models.
RECEIPTS: Transformer Circuitsfirst-party
May 21, 2024
ANTHROPIC MAPS MILLIONS OF FEATURES INSIDE CLAUDE
Anthropic researchers used sparse autoencoders to extract and map millions of concepts represented inside Claude 3 Sonnet, a major scale-up of Olah's interpretability program.
RECEIPTS: Anthropicfirst-party
Jul 6, 2026
THE LAB FINDS A 'GLOBAL WORKSPACE' PATTERN IN LANGUAGE MODELS
Anthropic published evidence for a global-workspace-like mechanism in language models, continuing the effort to explain model computation through identifiable internal structures.
RECEIPTS: Anthropicfirst-party