The video opens with a puzzle: your brain runs seeing, hearing, remembering, and speaking on completely separate hardware — yet you experience one single unified thought. "How do a crowd of specialists become one mind?" An intelligent system, the answer goes, isn't one giant blob. It's a collection of specialist modules — a vision module turns pixels into meaning, a language module handles words, a memory module recalls the past, a tool module calls out to the world. Each is its own little neural network, quietly producing its own output.
And that's where the trouble begins. Every module "speaks into the void," producing separate answers that never meet. The vision module sees a car; the memory module recalls a rule; the planning module wants to move — but nothing pulls those threads into a single coherent decision. If they never share, you don't get intelligence: "you get a room full of experts all talking at once and no one listening." This is the integration problem, and it's the thing the paper sets out to solve.
Cognitive science's leading answer is global workspace theory. Picture a theater: the specialist modules sit in the audience, each holding a piece of information. Above them is a stage with a single spotlight. At any moment, only one piece of information gets pulled into the spotlight and onto the stage. "That's the bottleneck, and the bottleneck is the whole point" — once something is on stage, it's broadcast back to the entire audience, every module sees the same thing, and the whole system updates together.
Anthropic's move is to build this inside an LLM. The shared stage becomes a common vector space — the J-space (J for Jacobian). The key rule: everything posted to it must be a vector of the same shape, a list of numbers in one common format. A model's internal thoughts live as points in a high-dimensional cloud of activations — but, crucially, knowing where a thought sits tells you almost nothing about cause and effect.
Location isn't enough. The question that matters for interpretability is: if I reach into that hidden state and nudge this one point a little, does the model's answer swing wildly, drift slightly, or stay completely still — and in which direction? To answer that you need an instrument that reads how the output responds to changes inside. That instrument is the J-Lens.
The visualization: picture the dim cloud of J-space with one activation glowing. Slide a lens in front of it, and suddenly arrows fan out from that point. Each arrow is a possible nudge — a direction you could push the hidden state. The brightness and length of each arrow shows how strongly the model's output reacts to that particular push. "The lens doesn't show you the point. It shows you the consequences of moving the point." That's the core reframe: interpretability isn't about reading activations directly, it's about measuring sensitivity.
The math starts with the simplest idea there is: slope. Take a curve, pick a spot, draw the tangent line. Its slope answers one clean question — for a tiny change in input, how big is the change in output? A steep slope means the output is sensitive here; a flat slope means you can wiggle the input and the output barely notices. That single number is the seed of everything.
Now the twist: inside a network there isn't one input axis, there are thousands — one for every dimension of the hidden state. Each axis gets its own tiny tangent, its own local slope. Collect every one of those slopes into a single grid and you have the Jacobian: "just a table" where every input direction is paired with its effect on every output. Feed it a nudge and it predicts the resulting change in the output. It's the multi-directional slope of the model, evaluated right where you're standing.
Here's the fact that makes interpretability possible: the directions are not equal. Shoot arrows outward from the marker in every direction and color them by steepness. Most directions come back cool and short — nudge that way and the output shrugs. But a handful glow hot and long: push along one of these and the output swings hard. "These few steep directions are the ones the model actually cares about."
Think of a mixing board with hundreds of sliders where most are wired to nothing. The lens tells you which few sliders are actually connected to the speakers. Then the payoff: rank the directions by how strongly they move the output, and look at what changes when you push each of the top ones. Push along the first steep arrow and the output's sentiment lifts; along the second, the tone grows more formal; along the third, the model's confidence rises. Attach human labels to those arrows, and the sea of flat, unlabeled directions simply fades away. Interpretability quietly becomes geometry.
Once you trust a direction, you can do more than read it — you can drive with it. Grab the sentiment arrow and push the glowing point along it, and the generated sentence shifts from grim to upbeat in real time. That's steering: editing behavior by adding a feature direction back into the activations. This is the part with real control implications — researchers can nudge specific behaviors, not just observe them.
But there are two caveats that keep it honest. First, superposition: "a labeled direction rarely means one clean concept." Concepts smear across directions, and directions mix concepts — the labels are useful, not exact. Second, locality: the lens only describes the neighborhood around one point. Drag the marker to a new spot and the whole fan rearranges — directions that were steep flatten out, new ones flare up. The sensitivity map is a snapshot, valid only where you took it.