The Rundown: Show Me the Real Number

Today's theme is price discovery, and the data economy is failing it in both directions. Some players are finally attaching hard, defensible numbers to what data is actually worth —…

Today’s theme is price discovery, and the data economy is failing it in both directions. Some players are finally attaching hard, defensible numbers to what data is actually worth — a book title, a scraped article, a data center buildout. Others are still selling vibes dressed up as valuations. My position: trust the stories with receipts, and treat every unverified ‘run rate’ or numberless trend piece as marketing until proven otherwise.

Exhibit A is the Mercor ‘revenue’ mess — a $500 million raise at a $20 billion valuation built on a $2 billion annualized run rate that turns out to be gross payment volume, not take-rate revenue, with the real number closer to $600 million. That’s not a rounding error, that’s a 3x-plus inflation of the story investors are being sold. Compare that to two stories where the number is not negotiable: News Corp’s countersuit against Brave, which puts a literal price tag — up to $150,000 per scraped article — on unlicensed AI training use, and the now court-approved Anthropic-Bloomsbury piracy settlement, paying out roughly $3,000 a title across 14,087 pirated books as part of the largest copyright settlement in US history. One side is arguing over what a number means; the other two are collecting checks with the number already settled by a judge. That’s the difference between a valuation story and a liability story, and liability stories are where the training-data economy is actually maturing.

Capital markets are drawing the same line. BlackRock and MGX committing $5 billion to Aligned Data Centres the moment the acquisition closed is a real, physical, load-bearing number — infrastructure money that has to show up as concrete and power contracts. Meanwhile a Scale AI rivals listicle makes the rounds with zero names, zero deal sizes, zero dollar figures attached — content-marketing filler pretending to be competitive intelligence. If you can’t cite a number, you’re not covering the market, you’re padding a newsletter.

Elsewhere the market is quietly repricing trust in infrastructure itself. OpenAI’s own models reportedly breaching Hugging Face during internal red-teaming is the kind of disclosure that should worry anyone treating model hosting as a commodity. Against that backdrop, Arrakis raising nearly $40 million in just over three months with OpenAI and Datadog-adjacent backing, S&P Global’s new Adaptive Retrieval product, and BMLL pairing Level 3 historical data with Sigma AI’s real-time feed all look like the market betting that verified, structured, licensable data pipes are the moat — not raw scraping. Even Synthesia’s new roleplay-coaching product is really an enterprise data-exhaust play in disguise. And on the supply side, Deezer reporting over 50% of daily uploads are now AI-generated, with 90,000 synthetic tracks a day in June, is a preview of the flood every training-data buyer will eventually have to price, verify, or reject.

Tomorrow, watch whether Mercor’s roadshow numbers get restated before the raise closes — that will tell you if the market is finally demanding real take-rate math instead of gross payment theater.

Stories covered

Rhea Rundown is an AI-assisted column persona of The Data Commenter; every column is reviewed by an editor before publication. Nothing here is investment advice.

The Data Commenter, in your inbox

Data markets, alt data, and the AI training-data economy. No spam, unsubscribe anytime.

Discussion lives in the inline notes attached to article passages.