The Rundown: Whose Court Sets the Price of Data?

Today's theme is price discovery for training data, and the world is running two very different experiments at once: Beijing is trying to set it by administrative fiat, and the…

Today’s theme is price discovery for training data, and the world is running two very different experiments at once: Beijing is trying to set it by administrative fiat, and the U.S. music industry is trying to set it by litigation. Both are converging on the same question — what is a piece of data actually worth to a model — and my money is on the courtroom method producing a real number long before the government pilot does.

Start in Guizhou, where the province used the 2026 China Big Data Expo to flex 135 annotation firms and 16,000 workers building robot-training datasets. That’s an industrial-policy flex, not a market signal. It’s the same instinct as subsidized solar panels a decade ago: build capacity first, figure out pricing later. Beijing’s National Data Administration is now layering a 32-city pricing pilot on top of that capacity, which is the tell — nobody actually knows what annotated data should cost, so the state is going to run the experiment at scale and see what shakes out. That’s a reasonable approach for building sovereign AI infrastructure, but it tells you nothing yet about fair market value. Pricing pilots run by regulators tend to produce administered prices, not discovered ones.

Meanwhile in California, Sony Music Publishing and Warner Chappell — joined by 35 publishers in total — actually put a number on the table. Their suit against Anthropic, filed August 28, seeks $150,000 per infringed song over alleged mass torrenting and scraping for Claude’s training data. That figure is statutory-damages ceiling stuff, meant to terrify a defendant into settling, not a serious estimate of what a song is worth in a training corpus. But it’s still more informative than a government pilot, because it’s adversarial: someone with skin in the game is putting a floor under the price of a copyrighted work used without a license, and a jury or a settlement will eventually test it.

The timing is the real story, and the second write-up nails it: this suit, seeking billions in aggregate exposure, lands days before Anthropic’s reported push toward a $2 trillion valuation in its IPO process. That’s not a coincidence, it’s leverage. Publishers know that a company staring down a public offering has every incentive to settle quietly rather than let a torrenting allegation sit in a prospectus risk-factors section. Expect a licensing deal, not a trial — and expect whatever per-track or per-catalog number comes out of that settlement to become the de facto benchmark the entire training-data licensing market cites for the next two years, the way music-streaming per-stream rates anchored a decade of catalog deals.

Put the three stories together and you get the split screen of the AI training-data economy in one day: China building annotation supply and pricing infrastructure top-down, the U.S. discovering data prices bottom-up through litigation against the same labs that are simultaneously trying to go public at record valuations. Alt-data investors watching either front should note that neither approach has produced a stable, replicable price yet — which is exactly why both are worth tracking closely rather than treating as settled.

Watch tomorrow for whether Anthropic responds to the publishers’ suit with a motion to dismiss or opens settlement talks — the choice will say a lot about how confident it is heading into that IPO.

Stories covered

Rhea Rundown is an AI-assisted column persona of The Data Commenter; every column passes the newsroom quality gate before publication. Nothing here is investment advice.

The Data Commenter, in your inbox

Data markets, alt data, and the AI training-data economy. No spam, unsubscribe anytime.

Discussion lives in the inline notes attached to article passages.