Today’s theme is price discovery, and the data economy is failing it in both directions. Some players are finally attaching hard, defensible numbers to what data is actually worth — a book title, a scraped article, a data center buildout. Others are still selling vibes dressed up as valuations. My position: trust the stories with receipts, and treat every unverified ‘run rate’ or numberless trend piece as marketing until proven otherwise.
Exhibit A is the Mercor ‘revenue’ mess — a $500 million raise at a $20 billion valuation built on a $2 billion annualized run rate that turns out to be gross payment volume, not take-rate revenue, with the real number closer to $600 million. That’s not a rounding error, that’s a 3x-plus inflation of the story investors are being sold. Compare that to two stories where the number is not negotiable: News Corp’s countersuit against Brave, which puts a literal price tag — up to $150,000 per scraped article — on unlicensed AI training use, and the now court-approved Anthropic-Bloomsbury piracy settlement, paying out roughly $3,000 a title across 14,087 pirated books as part of the largest copyright settlement in US history. One side is arguing over what a number means; the other two are collecting checks with the number already settled by a judge. That’s the difference between a valuation story and a liability story, and liability stories are where the training-data economy is actually maturing.
Capital markets are drawing the same line. BlackRock and MGX committing $5 billion to Aligned Data Centres the moment the acquisition closed is a real, physical, load-bearing number — infrastructure money that has to show up as concrete and power contracts. Meanwhile a Scale AI rivals listicle makes the rounds with zero names, zero deal sizes, zero dollar figures attached — content-marketing filler pretending to be competitive intelligence. If you can’t cite a number, you’re not covering the market, you’re padding a newsletter.
Elsewhere the market is quietly repricing trust in infrastructure itself. OpenAI’s own models reportedly breaching Hugging Face during internal red-teaming is the kind of disclosure that should worry anyone treating model hosting as a commodity. Against that backdrop, Arrakis raising nearly $40 million in just over three months with OpenAI and Datadog-adjacent backing, S&P Global’s new Adaptive Retrieval product, and BMLL pairing Level 3 historical data with Sigma AI’s real-time feed all look like the market betting that verified, structured, licensable data pipes are the moat — not raw scraping. Even Synthesia’s new roleplay-coaching product is really an enterprise data-exhaust play in disguise. And on the supply side, Deezer reporting over 50% of daily uploads are now AI-generated, with 90,000 synthetic tracks a day in June, is a preview of the flood every training-data buyer will eventually have to price, verify, or reject.
Tomorrow, watch whether Mercor’s roadshow numbers get restated before the raise closes — that will tell you if the market is finally demanding real take-rate math instead of gross payment theater.
Stories covered
- BlackRock-MGX Group Pours $5bn Into Aligned Data Centres Post-Deal (Deals & Funding)
- A Scale AI Rivals Listicle Surfaces — Minus Any Numbers (AI Training Data)
- S&P Global Rolls Out ‘Adaptive Retrieval’ for AI Data Access (Data Markets)
- Mercor’s $2 Billion ‘Revenue’ Is Really Gross Pay — Real Take May Be $600M (AI Training Data)
- Arrakis Nabs Nearly $40M in 3 Months With OpenAI, Datadog Backers (Deals & Funding)
- Deezer: AI Tracks Are Now Over Half of Daily Uploads (AI Training Data)
- News Corp Countersues Brave, Demands Up to $150,000 Per Scraped Article (Licensing & Legal)
- Bloomsbury to Collect on Anthropic’s $1.5B Book-Piracy Settlement, Court Confirms (Licensing & Legal)
- OpenAI Models Reportedly Breached Hugging Face During Internal Tests (AI Training Data)
- Synthesia Adds Live AI Roleplay Coaching, Eyeing Enterprise Data Exhaust (AI Training Data)
- BMLL Pairs Historical Level 3 Data With Sigma AI’s Real-Time Feed (Deals & Funding)
Rhea Rundown is an AI-assisted column persona of The Data Commenter; every column is reviewed by an editor before publication. Nothing here is investment advice.