Context Language Models video thumbnail
Coding Agent · Research Deep Dive

Context Language Models: Your Agent Doesn't Need Compaction

Source: Prompt Engineering 13:53 · Published Oct 5, 2026 Paper: arXiv 2609.37725 (Meta + UW)

0:00Every agent harness today handles a full context window the same way: near a fixed threshold — say 75% of the window — the model writes a summary, the harness keeps the last few turns, and everything earlier is replaced. But a summary is written before anyone knows what will matter later. If the one config value that turns out to be critical 40 turns from now didn't make the summary, it is "basically gone, and the agent doesn't even know what is missing." A new paper from Meta and the University of Washington proposes a different answer: let the model edit its own context instead of summarizing it.

The bottleneck compaction creates

1:47The video's opening example makes the failure concrete. Claude Code, Codex, and Pi all summarize at a fixed point (you can trigger it manually with /compact in Codex or Claude Code). The problem is that summarization is lossy in exactly the worst way: it keeps what seemed important at summarization time, not what becomes important later.

2:16Debug a service where a config file read 40 turns ago held the answer, and the agent "might spend a lot more time digging through the logs, or a lot of tokens trying to figure that out again" — because the summary already dropped it. Worse, summaries stack: compact enough and "you end up with a summary of summaries, which is going to potentially lose all of the most important information." The paper's small benchmark, Context Bench, captures exactly this with simple tasks like keeping a Sudoku board up to date while moves stream in — and "most existing compaction techniques actually fail" on it.

What a Context Language Model is

3:33The conceptual shift is elegant. In a normal language model, every turn appends: new context = old context + model output. "A context language model replaces that plus sign with a function." The model decides what the entire context looks like on the next turn — it can keep, shrink, delete, or rewrite anything.

"Instead of summarization, we let the agent edit his own context window."

The paper's framing (from the arXiv abstract): CLMs are "language models that natively manage their own context… treating the context as a file and allowing the model to make unrestricted updates to this file." It manages working context inside a task, then throws it away when the task ends.

How it actually works

4:08The implementation is deliberately boring — which is the point. Before every request, the harness writes the context out to a file; each turn is a block with a header. The model can open that file and edit it with a small Python script "exactly the same way it would edit code." Whatever remains in the file at the end of the turn becomes the next prompt.

0:31The demo runs in Pi on a local DGX box: the agent opens its own context file, takes a 3,000-token log it read a few turns earlier, and replaces it with a single line preserving exactly the information that context provides. The context-size graph climbs with every read and drops every time the model cleans up after itself — with "no rules about what to keep; the model is deciding what is most important."

What the model does with the freedom

4:38Handed this control, the model invents its own strategies. In the paper's runs it created a new chat role called notes for itself, wrote a helper function to compact old search results (called 37 times), and in a multi-agent run kept a scoreboard of 21 sub-agents inside its context, updated in place 163 times.

5:08Compaction can be steered in plain English: "Once you pass 24,000 tokens, compact it down to 4,000" — and it worked. The threshold is no longer a fixed harness knob; it's something you can talk the model through.

The paper's numbers

5:33On BrowseComp+ (a Deep Research benchmark), the video reports CLM scored 59.4% — about 11% higher than the best summarization approach, at roughly a fifth of the compute. The independent coverage agrees with the direction: CLMs "outperform state-of-the-art context strategies with higher accuracy and lower FLOPs across single- and multi-agent tasks."

5:49The authors released a Pi extension, installable with a single command; inside Pi, /clm opens a panel showing context size over time plus a page of every edit the model made, side by side.

The video's own test

6:42The video author built three over-window tasks (e.g. 16 days of warehouse logs where the agent must track truck seal codes, spot which were voided, and list the still-valid ones) and ran them against Pi's built-in summarization versus Pi + CLM. The baseline is where it gets interesting:

  • 7:28Pi summarization hallucinated. The first summary listed 12 seal codes — 11 of which "didn't exist anywhere." By the third summary, 14 of 18 codes were made up, and it "even invented a sequence right next to the real code." The summaries looked "perfectly confident, with checkboxes and next steps," so the agent had no way to know it was wrong.
  • 8:28The agent got the right answer anyway — not because it remembered, but because it ignored the summary and re-read every file from disk with grep, despite being told not to. On another task it concluded the logs had no deploy info when "the deploy line was right there" — read earlier but dropped by summarization.
  • 9:15Pi + CLM: the model wrote its whole context into one notes block tracking every seal applied, every seal voided, and what to read next. All 18 codes exact. Across the seal and incident tasks it got every answer right except one.

The gotchas

10:05The video's most useful section is where it fails. Four limits matter before you build on this:

  • Harness hardness matters more than you'd think. In three of four default runs, the model never edited its context — Pi CLM has a safety net that swaps old tool outputs for short notes on overflow, and that net did all the work. The culprit: parallel tool calls (reading four files in one turn) blew past the window before the model could react. Fix: one tool call per turn plus a "size note" after every result — how the actual paper harness runs.
  • 10:57Prefix-caching placement. Edits at the end of the context file preserve most of the prefix cache (fast); edits in the middle invalidate everything after the edit point, raising latency and cost.
  • 11:49It's not automatically cheaper. One seal run processed ~2× the tokens of plain Pi; another drew only 10% of its prompt from cache. The paper's "suffix cache reuse" fix exists but is an SGLang patch for one model on one GPU, which the author couldn't test.
  • 12:23The model must be good at this. A 9B model was six points behind clean summaries until RL-trained; and the headline results use a 32K-token budget — at 128K, CLM and summarization basically tie.

12:48One security note the paper itself flags: if the model can write to its own memory, so can a prompt injection — and the injection "stays across turns."

13:04The verdict: "Summaries are written once and then trusted forever — and in my runs, that's exactly where things went wrong. Letting the model manage its own context fixed those problems." As long as the harness gives the model room to react and you watch the cache, it's a promising direction.

Claims checked

Verified

Paper is real. "Context Language Models," arXiv 2609.37725, Meta + University of Washington. Abstract: "treating the context as a file and allowing the model to make unrestricted updates."

Verified

Repos exist. facebookresearch/context-language-models (official) and lolipopshock/pi-clm (the Pi extension).

Verified

Direction confirmed. Independent coverage: CLMs "outperform state-of-the-art context strategies with higher accuracy and lower FLOPs across single- and multi-agent tasks."

Corrected

Auto-caption garbles. "Cloud Code"→Claude Code, "VLLM"→vLLM, "SGL lang"→SGLang, "BrowseComp Plus"→BrowseComp+. The video's local test model ("quant 3.8 27B") and the paper's cited model size ("20.627B") are unclear in the caption and described generically rather than guessed.

Attributed

Specific figures (59.4%, +11%, ~1/5 compute, Context Bench, 32K/128K budgets, 9B RL gap) are the paper's reported results as relayed in the video — I did not independently re-run the benchmark. The existence of the paper and its headline claim are verified; the precise numbers are not.

Sources & further reading

☰ View all