AI NEWS PAGE DOSSIER

Chris Olah

Chris Olah's file: neural-network visualization, mechanistic interpretability, Anthropic, and the attempt to understand what giant models are doing inside.

Olah spent years inventing ways to see inside neural networks, moved from Google Brain and Distill into OpenAI's circuits work, then led transformer-circuits and large-model interpretability research at Anthropic.

HOW THE HELL DID CHRIS OLAH GET HERE?

  1. Sep 8, 2016

    HE STARTS DRAWING THE MACHINE'S INTERNAL LOGIC

    At Google Brain, Olah published a visual exploration of attention and augmented recurrent neural networks, part of a broader effort to make neural systems understandable by inspection.

    RECEIPTS: Distillfirst-party

  2. Nov 7, 2017

    HE TURNS NEURONS INTO PICTURES PEOPLE CAN READ

    Olah and collaborators published 'Feature Visualization,' a widely used framework for showing what image-network units respond to across layers.

    RECEIPTS: Distillfirst-party

  3. Apr 14, 2020

    OPENAI PUTS THE CIRCUITS UNDER A MICROSCOPE

    Olah contributed to OpenAI Microscope, a collection of visualizations meant to let researchers inspect the internal features and circuits of vision models.

    RECEIPTS: OpenAIfirst-party

  4. 2021

    INTERPRETABILITY MOVES INTO TRANSFORMERS

    Olah and collaborators published a mathematical framework for transformer circuits, extending mechanistic-interpretability methods from vision systems into attention-based language models.

    RECEIPTS: Transformer Circuitsfirst-party

  5. May 21, 2024

    ANTHROPIC MAPS MILLIONS OF FEATURES INSIDE CLAUDE

    Anthropic researchers used sparse autoencoders to extract and map millions of concepts represented inside Claude 3 Sonnet, a major scale-up of Olah's interpretability program.

    RECEIPTS: Anthropicfirst-party

  6. Jul 6, 2026

    THE LAB FINDS A 'GLOBAL WORKSPACE' PATTERN IN LANGUAGE MODELS

    Anthropic published evidence for a global-workspace-like mechanism in language models, continuing the effort to explain model computation through identifiable internal structures.

    RECEIPTS: Anthropicfirst-party

MORE DOSSIERS