Deep Dive · Interpretability

J-Space: Anthropic's Jacobian Lens for Reading LLM Thought

✍️ Arivu ⏱️ 6 min 📅 September 2026
Video thumbnail

1.The Integration Problem: Specialists That Never Meet0:00

The video opens with a puzzle: your brain runs seeing, hearing, remembering, and speaking on completely separate hardware — yet you experience one single unified thought. "How do a crowd of specialists become one mind?" An intelligent system, the answer goes, isn't one giant blob. It's a collection of specialist modules — a vision module turns pixels into meaning, a language module handles words, a memory module recalls the past, a tool module calls out to the world. Each is its own little neural network, quietly producing its own output.

And that's where the trouble begins. Every module "speaks into the void," producing separate answers that never meet. The vision module sees a car; the memory module recalls a rule; the planning module wants to move — but nothing pulls those threads into a single coherent decision. If they never share, you don't get intelligence: "you get a room full of experts all talking at once and no one listening." This is the integration problem, and it's the thing the paper sets out to solve.

2.Global Workspace Theory, Rebuilt as the J-Space0:19

Cognitive science's leading answer is global workspace theory. Picture a theater: the specialist modules sit in the audience, each holding a piece of information. Above them is a stage with a single spotlight. At any moment, only one piece of information gets pulled into the spotlight and onto the stage. "That's the bottleneck, and the bottleneck is the whole point" — once something is on stage, it's broadcast back to the entire audience, every module sees the same thing, and the whole system updates together.

Anthropic's move is to build this inside an LLM. The shared stage becomes a common vector space — the J-space (J for Jacobian). The key rule: everything posted to it must be a vector of the same shape, a list of numbers in one common format. A model's internal thoughts live as points in a high-dimensional cloud of activations — but, crucially, knowing where a thought sits tells you almost nothing about cause and effect.

The paper, in one line: "Verbalizable Representations Form a Global Workspace in Language Models" (arXiv 2607.15495) — the idea that a small, shared activation space holds the concepts the model is actively reasoning with, and that you can read it through the model's own vocabulary. Companion code: github.com/anthropics/jacobian-lens (Python, Apache-2.0).

3.The J-Lens: Reading Consequences, Not Positions1:53

Location isn't enough. The question that matters for interpretability is: if I reach into that hidden state and nudge this one point a little, does the model's answer swing wildly, drift slightly, or stay completely still — and in which direction? To answer that you need an instrument that reads how the output responds to changes inside. That instrument is the J-Lens.

The visualization: picture the dim cloud of J-space with one activation glowing. Slide a lens in front of it, and suddenly arrows fan out from that point. Each arrow is a possible nudge — a direction you could push the hidden state. The brightness and length of each arrow shows how strongly the model's output reacts to that particular push. "The lens doesn't show you the point. It shows you the consequences of moving the point." That's the core reframe: interpretability isn't about reading activations directly, it's about measuring sensitivity.

4.The Jacobian: Slope, Gridded into Every Direction2:45

The math starts with the simplest idea there is: slope. Take a curve, pick a spot, draw the tangent line. Its slope answers one clean question — for a tiny change in input, how big is the change in output? A steep slope means the output is sensitive here; a flat slope means you can wiggle the input and the output barely notices. That single number is the seed of everything.

Now the twist: inside a network there isn't one input axis, there are thousands — one for every dimension of the hidden state. Each axis gets its own tiny tangent, its own local slope. Collect every one of those slopes into a single grid and you have the Jacobian: "just a table" where every input direction is paired with its effect on every output. Feed it a nudge and it predicts the resulting change in the output. It's the multi-directional slope of the model, evaluated right where you're standing.

5.Steep Directions and the Mixing Board4:30

Here's the fact that makes interpretability possible: the directions are not equal. Shoot arrows outward from the marker in every direction and color them by steepness. Most directions come back cool and short — nudge that way and the output shrugs. But a handful glow hot and long: push along one of these and the output swings hard. "These few steep directions are the ones the model actually cares about."

Think of a mixing board with hundreds of sliders where most are wired to nothing. The lens tells you which few sliders are actually connected to the speakers. Then the payoff: rank the directions by how strongly they move the output, and look at what changes when you push each of the top ones. Push along the first steep arrow and the output's sentiment lifts; along the second, the tone grows more formal; along the third, the model's confidence rises. Attach human labels to those arrows, and the sea of flat, unlabeled directions simply fades away. Interpretability quietly becomes geometry.

What the lens finds in practice: the J-space holds things the model knows before it writes a word — hidden bugs in code, the intermediate steps of a math problem, suspicion that a prompt is an injection attempt. Switch off the J-space and reasoning, analogy, and translation break down, while simple recall stays mostly intact. And chain-of-thought? It works like scratch paper for the model — an externalized version of exactly this workspace.

6.Steering, Superposition, and the Honest Limits5:04

Once you trust a direction, you can do more than read it — you can drive with it. Grab the sentiment arrow and push the glowing point along it, and the generated sentence shifts from grim to upbeat in real time. That's steering: editing behavior by adding a feature direction back into the activations. This is the part with real control implications — researchers can nudge specific behaviors, not just observe them.

But there are two caveats that keep it honest. First, superposition: "a labeled direction rarely means one clean concept." Concepts smear across directions, and directions mix concepts — the labels are useful, not exact. Second, locality: the lens only describes the neighborhood around one point. Drag the marker to a new spot and the whole fan rearranges — directions that were steep flatten out, new ones flare up. The sensitivity map is a snapshot, valid only where you took it.

What it does — and doesn't — say about consciousness: finding a global-workspace-like structure in an LLM is suggestive, but the video is careful. It shows a bottleneck-and-broadcast pattern in the activations, not sentience. The honest read: a shared workspace is a functional requirement for integrated behavior, and LLMs appear to implement something like it — which is an interpretability result, not a claim about subjective experience.

Key Takeaways

  1. The integration problem: specialist modules (vision, language, memory, tools) produce separate outputs that never meet — intelligence requires pulling them together.
  2. Global workspace theory solves it with a bottleneck-and-broadcast stage; Anthropic rebuilds that as the J-space, a shared vector space of same-shape activations.
  3. The J-Lens doesn't read positions — it measures sensitivity: how strongly the output reacts to a nudge along each direction.
  4. The Jacobian is the multi-directional slope — a grid pairing every input direction with its effect on every output.
  5. Most directions are flat (output shrugs); a handful are steep — those are the concepts the model actually cares about, and they carry human labels.
  6. Steering = pushing a labeled direction to change behavior in real time (grim → upbeat).
  7. Two honest limits: superposition (concepts smear across directions) and locality (the sensitivity map is valid only where you took it).
  8. The J-space holds bugs, math steps, and injection suspicion before the model writes a word — and CoT works like scratch paper for it.

Timestamp Index

☰ View all