Is Anthropic stealing your data?

You Pay Twice: Aggregate Usage Data and the AI Sovereignty Ladder

The headline question gets a one-word answer, and it is not the interesting part. What follows is a reading of the actual terms of service, a four-case pattern of platforms entering their customers' verticals, and a six-tier ladder for deciding how much of your stack you actually own.

Video thumbnail
📺 IndyDevDan ⏱️ 34:29 📅 27 July 2026
AI sovereignty Terms of service Open weights IP defence Agentic engineering Platform strategy

💰 The second payment 0:00

The framing device comes from Satya Nadella, and the host quotes it directly: you pay for intelligence twice — once with money, and again with the proprietary knowledge you must reveal to make that intelligence useful. The better you want the model to perform, the more of that knowledge you have to feed it.

Alex Karp of Palantir supplies the second quote, about technical customers wanting control over their compute, their models, their data stack, their alpha — and wanting to know the means of production is not being transferred to someone else.

The host is explicit that he does not agree with either take entirely, but considers both valid. He also sets a standard for the episode that is worth holding him to: if the question cannot be answered with supporting ground truth data, it does not get answered at all.

His own framing is that language models are meaningfully worse than search on this axis. A search engine gets your query. An agent gets the workflow, the traces, the system prompts — the operating knowledge of the business.

⚖️ The answer, up front 1:58

The headline question gets answered in the first third of the video, which is the right structural choice: no, Anthropic is not stealing your data.

The verdict: They are not stealing your data. They are using it anonymized, in aggregate — and that aggregate steers what they build next. The uncomfortable part is not a violation. It is what the terms permit.

The host repeats throughout that this is not an attack on Anthropic, and that he is a genuine fan. He also generalises the argument: this is a story about what platforms are, and Google and OpenAI run the same play. Anthropic is simply the largest and clearest example to point at.

🌀 The tumbler 3:16

The central metaphor is a coin tumbler. Every customer's prompts go into the same bucket, get anonymized and randomised, and what comes out the other side is no longer attributable to anyone.

The host's point is that anonymisation does not destroy the signal it is usually assumed to destroy. You can still identify the value proposition each prompt, each agent, each trace delivered. The output is not your data — it is a map of the market.

Why this is fair play, and still a problem: the anonymisation is real, disclosed, and honoured. The aggregate that results is nobody's property in particular, which is precisely why it can be acted on freely.

He also draws the ownership line sharply: when you prompt in Claude Code or Codex, you are renting the model. The harness, the traces, the prompts, the system prompts, the outputs — all yours. The intelligence underneath is not. From fifteen years of engineering he names it in familiar terms: dependency risk, keyman risk, key technology risk.

🔁 One is coincidence, four is confirmation 4:20

This is the evidential core of the argument, and it is presented as a sequence rather than an accusation:

Usage signalWhat followed
Cursor scalesClaude Code
Figma MCP usage explodesClaude Design
Security tooling takes offClaude Security
Life science usageClaude Life Science

The host's own formulation: one is a coincidence, three is a pattern, four is a confirmation of the pattern. He grants that competing in a single domain would be entirely reasonable for a platform with that access — it is the repetition that changes the reading.

He describes the mechanism as a four-step process: see the usage → learn the trends → enter the vertical → cut access. And then, importantly, he qualifies his own claim. The access-cutting largely stopped; he cites the earlier restrictions on Cursor and Windsurf as historical rather than current. The principle survives even though the practice has faded.

📄 Clio, and what the docs say 6:20

The host is emphatic that none of this is speculation or leaked information — it is in the terms of service, and users agreed to it.

Clio is Anthropic's aggregated, privacy-preserving analysis system for gaining insights into real-world impacts and usage patterns while maintaining user privacy. The host's reaction is genuinely positive: without such a system, many enterprises simply could not use the technology at all.

He then connects it to something visible. The Anthropic Economic Index — the published breakdowns of how AI is used across regions, sectors and socio-economic groups — is built from this data. There is no secret about it.

The setting worth changing

A practical aside, demonstrated live: in Claude, Cmd+Shift+, opens settings, and under Privacy sits the help improve our AI models toggle. The host's recommendation is to switch it off if you are doing IP-sensitive work, and notes that the privacy statement's "may conduct aggregated anonymized analysis" is, in practice, "does".

He gives credit where it is due: the data is not sold to third parties, and deletion requests are honoured promptly. He also flags one carve-out — a 30-day retention requirement tied to cybersecurity harms, the same policy that reportedly triggered a ban in one government context.

🎚️ Not all customers are equal 8:19

The tiering is straightforward once stated, and most users have never checked which tier they are in:

TierData treatment
FreeEverything is used, including for training. If it is free, you are the product.
Pro / MaxConsumer terms. Data feeds Clio and the aggregate reports.
Teams / Enterprise / APICommercial terms. Materially stronger protections.
The inheritance rule: Claude Code, Claude co-work, and anything else you use inherits the protections of the account you are logged into. The tool does not determine your privacy posture — your contract does.

This is the hinge between the problem and the solution: the same keystrokes carry different data protections depending entirely on which account issued them.

🔒 Incentives, not goodwill 9:56

The host's answer to "why should we believe them?" is refreshingly unsentimental. It is not trust, and it is not corporate values.

"If they violate a term of service even one time, that enterprise is going to discover there's going to be a mass enterprise exodus and then the company is cooked. It's incentives that protect your and I's data, not goodwill."

He adds the honest limit of his own analysis: he is addressing what is in the public record. What happens outside it, nobody knows — and he means that for every lab, not one of them.

✅ Four claims, two survive 10:44

The host invites viewers to guess which of four claims hold up before revealing the answers:

ClaimVerdictBasis
They see aggregate patternsTRUEDisclosed; this is what Clio does
They train on your code and promptsFALSETerms of service say no
They own your outputsFALSETerms of service say you own them
They compete with your productTRUEFour verticals entered, and counting

The host is careful about why the fourth is not a loophole: it is aggregate data, not your data. Everyone's, not yours. That distinction is what makes it permissible, and it is written into the agreement rather than hidden from it.

The test he offers is a single question: are you in a growing domain, and how big? If the answer is a large growing domain, expect a lab to enter it eventually. He also notes why this is more plausible now than it was — what a single engineer can accomplish with agentic tooling is exponentially higher than a few years ago, which makes entering a vertical far cheaper than it used to be.

🧪 Commodity agents vs IP agents 14:14

The solution starts by refusing to solve the problem where it does not exist. The host frames it as a senior engineer's instinct: if you do not have to solve the problem, do not waste time on it.

Commodity agentsIP agents
ContentsPrototypes, CRUD, boilerplate, glue codeDomain logic, hand-farmed evals, traces, workflows, user-data insights
CharacterReplaceable; anyone could prompt it in a weekendScarce, asymmetric, compounding
DistributionInside the normal distribution the models already knowThe sub-percent tail the training data never saw
ActionShip it; do not lose sleepDefend it
The one-sentence test: if a competitor could read your full agent trace, would it matter? The host says that if you remember nothing else from the video, remember this — it is the 80/20 of the whole solution.

He is blunt that most engineering work is commodity work, and says so without condescension: that work is still valuable, it is simply not IP-defensible. By his estimate the genuinely defensible portion is 20%, or 10%, or 5%. For most engineers, the rest of the video does not apply — and he says so directly rather than manufacturing urgency.

🪜 The AI sovereignty ladder 17:41

The ladder is the practical spine of the episode. Each rung trades convenience for control:

TierWhat it isData exposure
0Consumer subscriptionLowest protection; used in aggregate
1Commercial APIBetter terms; only sampled in aggregate, ZDR options
2Own the control plane — your VM, gateway, routerYou hold every trace; switch models freely
3Model cloud (AWS, GCP Vertex, Azure Foundry)No lab sampling at all
4Hybrid private — rent GPUs, own the open-weight modelOnly GPU usage leaves your control
5On-prem metalTotal, and financially out of reach for almost everyone

Two moments of intellectual honesty stand out here. First, the host corrects himself mid-explanation: tiers 2 and 3 are probably in the wrong order, because an LLM gateway and a VM give you less protection than moving onto a cloud provider. He says plainly that he made a mistake and needs to swap them.

Second, he is candid about the ceiling. Owning GPUs on premises is the theoretical optimum and, in his words, financial suicide unless you are very wealthy or can write serious cheques. His recommended landing zone is tier 4: rent the GPUs, own the model.

The cascade argument for tier 4: at tiers 0–3, an external decision can remove your access — including a government ruling that a model may not be served. At tier 4 you hold the weights, so the decision is no longer someone else's to make.

🎯 Which tier is yours 19:38

The recommendations are segmented by who you are, not by maximalism:

Who you areRecommended tierWhy
Individual engineerTier 0–1Best return on both IP protection and spend; commercial API only samples in aggregate
Individual, $1M+ revenue, non-commodity workMove upThe economics now justify the effort
Small / medium business (2–100)Tier 2–3On a model cloud your data is not sampled at all
EnterpriseTier 4Own the model, the traces, the evals, the routing

On the model cloud option he explains the mechanism: Anthropic has arrangements with the major clouds to mount its technology on their hardware and then step away. You transact with your cloud provider, not with the lab.

He also addresses the objection he clearly received on a previous video about API spend, and does not retreat from it: compute is still relatively expensive, but the price per intelligent agent task keeps falling. Spend to win.

🔓 Open weights, with a caveat 23:41

Open-weight models are named as the only real escape hatch — Kimi K3, GLM 5.2, Minimax, the Qwen family. The host gives the Chinese labs genuine credit: they are scorching the earth and commoditising the closed US labs, and he frames that as good for engineers.

The distinction he insists on: open weights does not mean overseas discount APIs. Downloading and running the model is sovereignty. Calling someone else's endpoint is the same dependency in different clothes.

His reasoning is about enforceability rather than accusation: terms of service across those jurisdictions cannot be reliably enforced. He acknowledges his own position — he says he is skeptical partly because he is a US engineer, and concedes he cannot prove US labs are cleaner, only that he does not believe otherwise. He is empathetic toward anyone early in their career for whom the discount APIs are the only affordable option.

One more honest caveat: self-hosting wins on secrecy and trace ownership, but not automatically on price. Moving up the ladder requires means, and means come from success — cash and product-market fit. If you have no business and are simply paranoid, he says, this will be cost-prohibitive.

🔮 Three scenarios, one defence 25:52

The closing analysis lays out three futures and marks clearly which is speculation:

ScenarioWhat happensHost's read
BestLabs stay labs — best models, best prices, out of the application layerWould love it; only NVIDIA looks close
MidCurrent behaviour continues; platforms keep entering verticalsWhat he is betting on
WorstAnonymisation is a catch-all and prompts get absorbed into lab IPExplicitly flagged as speculation he does not believe

The worst case is where the host is most careful. He describes a possible "legal alchemy" — once data is anonymized it is no longer yours, so it can be trained on or built against — and immediately labels it as ungrounded speculation he does not buy, while noting it is probably close to what Alex Karp believes.

The point of the exercise: the defence is identical in all three scenarios. Own the traces, own the evals, keep a second model path for IP-essential work, rent the GPUs and own the model. You do not need to resolve which future is real to know what to do.

The final answer restates the verdict. Anthropic is not stealing your data — the public record says so and the incentive structure backs it. But your usage shapes what they build next, and the list of verticals is easy to extend. Same verdict for OpenAI. This is what platforms are.

The closing image is the one from the opening: never has a technology asked you to pay in cash and in know-how. The host's answer is not to stop using these tools — he is emphatic about admiring the work — but to own the learning loop and run your own router, so that access to the frontier is a choice rather than a dependency.

💡 Key takeaways

  1. The answer to the title is no. They are not stealing data and not training on your prompts — the terms of service say so, and the incentive structure enforces it. The interesting question is the one underneath.
  2. Anonymised aggregate is still a market map. Tumbling everyone's traces together removes attribution but preserves the signal about which problems are valuable and growing.
  3. Four verticals is not a coincidence. Cursor → Claude Code, Figma MCP → Claude Design, security tooling → Claude Security, then Life Science. The host's rule: one is coincidence, three is a pattern, four confirms it.
  4. None of this is leaked or hidden. Clio is documented, the Economic Index is built from it, and users agreed to the terms. The discomfort is about what is permitted, not about a violation.
  5. Your privacy posture is your contract, not your tool. Free means you are the product; Pro and Max sit under consumer terms; teams, enterprise and the commercial API get materially stronger protections. Claude Code inherits whichever account you are in.
  6. Incentives protect you, not goodwill. A single terms violation would trigger an enterprise exodus. That commercial exposure is the real firewall.
  7. Ask one question about every prompt. If a competitor could read your full agent trace, would it matter? That single test separates commodity work from IP work.
  8. Most work is commodity work, and that is fine. The host puts genuinely defensible work at 5–20% and tells the rest of the audience the video does not apply to them.
  9. The ladder has six rungs. Consumer subscription → commercial API → own control plane → model cloud → rent GPUs and own open weights → on-prem metal. Climb as the business grows.
  10. Tier 4 is the recommended landing zone. Rent the GPUs, own the model. On-prem hardware is the theoretical optimum and, in his words, financial suicide for almost everyone.
  11. Open weights does not mean overseas APIs. Download and run the model, or you have swapped one dependency for a less enforceable one.
  12. Self-hosting wins on secrecy, not on price. Climbing the ladder requires means, and means come from revenue and product-market fit.
  13. The defence is the same in all three futures. Own the traces, own the evals, keep a second model path, rent GPUs and own the model. No forecast required.
  14. The host marks his own uncertainty. He corrects his tier ordering mid-argument, labels the worst-case scenario as speculation he does not believe, and concedes his skepticism of overseas labs is partly positional. That is what makes the grounded claims worth weighing.

🔗 Resources & links

🕐 Timestamp index

0:00Is your AI agent IP safe? — Nadella and Karp
1:58Is Anthropic stealing your data?
2:03Visualising the problem — paying twice
2:42Roadmap for the video
3:16No, but used anonymized — the tumbler
4:20The market map: coding, design, legal
4:26Four verticals — the pattern confirmed
5:06The four-step platform process
5:31You do not own the model — dependency risk
6:20Clio — aggregated privacy-preserving analysis
7:00The Economic Index is built from it
7:28The privacy setting to switch off
8:04The 30-day retention carve-out
8:19Not all customers are equal — consumer vs commercial
9:41Claude Code inherits your account tier
9:56Only incentives protect your data privacy
10:44Four claims — which two survive
11:25The verdict: not stealing, but competing
12:17The test: are you in a big growing domain?
13:04What can we do to defend our agentic IP
14:14Commodity agents vs IP agents
15:27The one-sentence test — the full agent trace
16:35Privacy is a stack, not one piece
17:41The AI sovereignty ladder
19:38Individuals — tiers 0 and 1
20:52Small and medium business — model cloud
21:26The self-correction: tiers 2 and 3 swapped
21:50Enterprise — own the model
22:41The cascade — why tier 4 survives external decisions
23:41Open-weight models are the future
24:04Do not confuse overseas APIs with owning the model
24:53Self-hosting wins on secrecy, not on price
25:52Best, mid and worst case scenarios
26:45The worst case, explicitly marked as speculation
27:09The defence is the same in all three
28:00Same verdict for OpenAI — this is what labs are
28:45Own the learning loop, run your own router
29:54Recap — the market map and the ground truth
32:12Do not hand off your IP without defences
33:22The double attack: capital and intellectual property