OpenRouter State of Models: Jev, Open-Weights, and Tokenomics

Five trends every engineer needs to understand entering Q4 2026 โ€” token usage that refuses to flatten, the "everyone is winning" race, a new species of decision model, the real cost of "free" tokens, and why you should own your agent harness.
Video thumbnail
IndyDevDan 26:08 Published Oct 5, 2026
LLM Market Tokenomics Jev OpenRouter Pi Agent Harness Engineering

The State of Models, Q4 2026

โ–ถ 0:00

With the release of the new Jev System One model class and the overnight wave of clones it spawned, IndyDevDan (Andy) uses OpenRouter's public data to break down where LLMs actually stand as we enter the final quarter of 2026. The thesis, stated up front: there is no AI bubble โ€” the evidence is the token-usage chart, and everything downstream hangs on whether it keeps climbing.

The five trends he walks through: tokens going up nearly exponentially, who's actually winning the AI race, the return of the zero-shot classifier in Jev, the hidden cost of "free" models, and the quiet rise of the Pi coding agent as a threat to Claude Code's position. Each maps to a concrete decision about which model to deploy at what performance, speed, and cost.

Trend 1: Tokens Go Up

โ–ถ 0:40

The headline number: roughly one year ago OpenRouter handled 4.5 trillion tokens in a week. Last week it was 146 trillion tokens. That's not exponential, Andy notes, but it's close โ€” and it's the single strongest argument against the bubble narrative, because it's concrete demand measured in the open, not a story.

The uncomfortable question he poses is about tokenomics, a word he uses deliberately. "Is your token spend going up because you're getting more useful work done, or because you're burning tokens to feel productive?" More tokens does not automatically mean more value. He draws the line at his own channel motto โ€” "scale your compute to scale your impact" โ€” and warns that burning tokens to feel productive is not progress. "Generating code is the easiest part of software engineering. The hard part is everything else."

The financial stakes sit squarely on Anthropic, which is targeting a $2 trillion IPO while losing almost $50 billion. If usage keeps climbing the way the OpenRouter chart does, Anthropic hits its $200 billion annualized-revenue projection by end of 2028 and the IPO is justified. If it stalls, "the house of cards starts falling down, and OpenAI likely follows." Andy is in the optimistic camp, but the bet is live and will be decided soon.

Trend 2: Who's Winning the AI Race?

โ–ถ 5:06

By OpenRouter market share, the top three have been DeepSeek, Google, and OpenAI for almost the entire year, with "others" occasionally poking into third. Andy pushes back hard on the "Google has fallen out of the race" take: the Gemini Flash series is his workhorse. Specifically, Gemini 3.8 Flash is the default model inside his Pi coding agent, chosen for the best performance/speed/cost trade-off.

But he insists the market-share chart must not be read at face value, because OpenRouter only counts tokens that flow through OpenRouter โ€” mostly individual engineers and SMBs, not enterprise. The napkin math: if Anthropic's rumored $60 billion run rate were all Opus 5.5 tokens, that's ~65 trillion tokens per day, or ~450 trillion a week. Cut it in half, cut it in half again โ€” it's still 1.5ร— OpenRouter's entire weekly volume through direct traffic. "The AI labs are getting massive, massive direct traffic."

His verdict on the race is deliberately anticlimactic: everyone is winning. Token spend is up 50โ€“80ร— across the board year-over-year. Chinese open-weight models make the technology affordable, frontier labs pave the way, and there's a "circular optimization" โ€” open-weight models increasingly being trained on distillations of the state of the art. Cost down, intelligence up.

Trend 3: Jev โ€” A New (Old) Model Class

โ–ถ 10:27

What's old is new again: Jev, by TypeSafe, brings zero-shot classifier models back โ€” and it's already 12th on OpenRouter despite using a fraction of the tokens an LLM burns inside an agent. OpenRouter has added a new "decisions" category specifically for this model class. Jev is also, per Andy's chart, remarkable for its uptime โ€” "Jev just doesn't go down," in contrast to the latency and reliability blips of the big labs (he films mid-outage as Claude stumbles again).

The claim worth scrutinizing is that everyone thinks they can "just clone Jev" with a fine-tuned scikit-learn model. Andy's counter: Jev is a generic zero-shot classifier applicable across many scenarios. If you have a specific dataset, yes, a small model can beat it on that one task โ€” but not across dozens of use cases, and not with Jev's reliability. "Contrary to popular belief, Jev is not easy to replace."

In practice: embedding Jev into his Pi agent as a routing/decision layer, Andy reports saving ~20% of token spend on a 400โ€“500K input prompt โ€” a small absolute number (5ยข), but the percentage compounds across every agent run.

The strategic point: Jev is a new species of model that sits alongside LLMs in an agent โ€” "Combine compute, don't select compute." It's not about any single model winning; it's about routing the right decision to the right model at the right time.

Trend 4: "Free Tokens"

โ–ถ 15:22

Stealth free models like SpaceBunny Alpha (rumored to be the next Minimax) rocket up the rankings purely because they cost nothing. Andy's warning is the classic one: if it's free, you're the product. The providers are almost certainly retaining your prompts and training on them, whatever the terms claim. "I'd rather pay than have them take proprietary information."

The bigger signal he extracts from free-model adoption: the cheaper effective models get, the more tokens get spent. Cheaper-but-capable models rise in market share, and low-cost launches drive outsized attention and usage on price alone. That's the trend pointing at the future โ€” and a reminder that OpenRouter's user base skews price-sensitive, which is exactly why its data over-indexes on cheap open-weight models and under-counts the enterprises paying top dollar for Claude and OpenAI.

Trend 5: Pi Agent and Harness Control

โ–ถ 19:23

On OpenRouter's top-apps list, the Pi coding agent (๐Ÿ”— pi.dev) has spent a long time trailing Claude Code in the fifth-to-seventh slots โ€” and it's been slowly, steadily closing the gap, while Codex trails behind. Andy's year-end prediction, teed up here early: Pi passes Claude Code in OpenRouter usage.

The read-through is about harness engineering. "The more engineers realize how important it is that they own their agentic coding tool โ€” that they can customize it, that they can control it โ€” the bigger this number gets." He notes the caveat that Claude Code's real numbers are far higher once you include Anthropic memberships and subscriptions, but the direction of travel is what matters.

The principle: "We already rent our intelligence. Don't rent your harness too." Own the thing you can control โ€” and by knowing how to harness-engineer (adding Jev tools, modifying the out-of-the-box experience), Andy saves ~20% token spend on every agent he boots.

He also flags the other open-weight harness options โ€” Cline as the best follow-up to Pi, Kilo Code as not bad โ€” but none match Pi's simplicity and customization.

Combine Compute, Don't Select It

โ–ถ 22:08

The through-line, repeated as his operating philosophy: "Combine compute. Don't select compute." It's DeepSeek V4.1 Flash and Jev; it's Claude Opus 5.5 and GPT-6 Astra together. The engineers who win understand the full menu of intelligence available to them and scale their compute to scale their impact.

His working method is concrete: throw state-of-the-art models at a new problem first, then scale down to cheaper, faster models once the work repeats. That's the practical core of "tokenomics" for agentic coding work โ€” and it's why he maintains a model-stack tier list and benchmarks long-horizon out-loop agentic coding.

On the bubble question he closes with a measured view: "There isn't really an AI bubble. This is very raw, useful technology." The real risk is financial, not technological โ€” how much the US economy is banking on the Anthropic and OpenAI IPOs going well. "That could pop the financial AI bubble." On the technology side, growth keeps defying every predicted wall.

Sources & Claims Checked

๐Ÿ“Ž Sources

This is a market-roundup video whose facts I spot-checked against primary sources before publishing:

Claims checked, with two flags:

  • โœ… Anthropic $2T IPO / ~$50B loss โ€” corroborated by Yahoo Finance and the Wall Street Journal's reporting.
  • โœ… Jev by TypeSafe as a zero-shot decision/classifier model โ€” confirmed by Wikipedia, Glean's enterprise evaluation, and independent benchmarks.
  • โš ๏ธ The headline token number. The video's description and main narration say 146 trillion tokens/week (up from 4.5T a year ago, โ‰ˆ32ร—). The closing summary slips to "445 trillion" and "80ร—" โ€” inconsistent with the chart and description. This article uses 146T and treats the growth as "โ‰ˆ30ร—," not 80ร—.
  • โš ๏ธ SpaceBunny Alpha / "next Minimax" โ€” explicitly a rumor in the video; treated as such here.

Transcription notes: the auto-captions rendered "Cline" as "Klein" and "GPT-6 Sol/Luna" as "6 soul / six luna"; corrected here silently against the known model names.

Key Takeaways

  1. Token usage is the bull case. OpenRouter went from ~4.5T to ~146T tokens/week in a year (โ‰ˆ30ร—); as long as it climbs, "there is no bubble."
  2. Tokenomics โ‰  more tokens. More tokens don't equal more value โ€” the question is whether spend buys useful work or just feels productive.
  3. Everyone is winning the race. DeepSeek, Google, and OpenAI hold the OpenRouter top three, but Anthropic's direct enterprise traffic may be 1.5ร— OpenRouter's entire volume.
  4. Google isn't out. Gemini 3.8 Flash is Andy's default model in Pi โ€” the best performance/speed/cost intersection.
  5. Jev is a new model species. TypeSafe's zero-shot decision model hit 12th on OpenRouter on a fraction of the tokens, and it's not trivially cloneable.
  6. Classifier models save real tokens. Embedding Jev as a router in an agent cut ~20% of token spend in Andy's benchmark.
  7. "Free" models cost you data. If it's free, you're the product โ€” free providers are almost certainly training on your prompts.
  8. Own your harness. Pi is closing on Claude Code in OpenRouter usage because engineers want to customize and control their agent tool โ€” don't rent the harness too.
  9. Combine compute, don't select it. DeepSeek V4.1 Flash + Jev + Claude Opus 5.5 + GPT-6 Astra together, routed at the right time.
  10. Throw SOTA first, then scale down. Use top models on new problems, cheaper/faster models on repeated work โ€” the practical core of agentic tokenomics.

Timestamp Index

0:00 State of Models
0:40 Trend 1: Tokens Go Up
5:06 Trend 2: Who's Winning
10:27 Trend 3: Jev
15:22 Trend 4: Free Tokens
19:23 Trend 5: Pi Agent
22:08 Combine Compute
โ˜ฐ View all