The Rundown: Who Actually Owns the Data You’re Buying?

The theme today is provenance, and the market keeps flunking the test. Every story in today's wrap is really the same story: who controls the underlying data, who profits from…

Who controls the underlying data, who profits from it, and whether the buyer or regulator can actually verify the chain of title. My position is simple — in 2026, provenance isn’t a compliance footnote, it’s the whole deal. Companies structuring around it, rather than through it, are building on sand.

Start with the cleanest example of structuring-around-it: NaaS Technology’s $15 million related-party EV data deal. A company buying data from itself, essentially, and calling it an arm’s-length agreement. Fifteen million dollars is real money, and related-party structures aren’t automatically fraudulent — but when the seller and buyer share ownership, the burden of proof on data quality and pricing shifts entirely to disclosure, and disclosure is exactly what’s thin here.

Then there’s the training-data reckoning finally arriving in court. Suno’s parallel copyright cases in Munich and Boston are the clearest test yet of whether “we scraped it and called it fair use” survives contact with two different legal systems in the same month. If Boston and Munich diverge — and there’s every reason to think they will, given how differently the U.S. and Germany treat text-and-data-mining exceptions — every AI company training on unlicensed catalogs needs to start pricing in geography-specific liability, not just aggregate risk. That’s a very different M&A calculus than the one most training-data deals were priced on eighteen months ago.

Meta’s own reversal on the Instagram @-mention image tool belongs in the same bucket, even though it’s product, not licensing. Ship a feature that lets anyone generate images of tagged public accounts, get instant backlash over nonconsensual likeness use, pull it days later. The pattern is consistent: consent gets treated as a launch obstacle to route around rather than a design constraint to build in, and it keeps blowing up in exactly the same way.

Against that mess, two stories look almost quaint in their orderliness. MDA Space’s acquisition of France’s CLS for its ocean and maritime monitoring data is a straightforward vertical buy — even if, notably, deal terms are absent from the announcement, so let’s not get too self-congratulatory about transparency in this sector either. And Salesforce and Databricks deepening their alliance around “trusted” enterprise AI data is the vaguest story of the day — heavy on the word “trusted,” light on what actually changed. Trust as a marketing word is cheap; trust as an audited data lineage is the thing enterprises are actually going to pay for.

That’s the throughline into the piece on data sovereignty becoming a strategic moat for agentic AI. As agentic systems chew through proprietary enterprise data, the question of where that data physically sits and who monetizes it stops being an IT question and becomes the entire competitive advantage. NaaS’s related-party structure, Suno’s scraping, Meta’s tagging tool — all of them are attempts to monetize data without fully settling who owns it. The companies that settle that question first, cleanly, are the ones whose moats will actually hold.

Watch tomorrow: any early signal from the Boston or Munich Suno rulings will move the whole training-data licensing market, so that’s the one to check first.

Stories covered

The Data Commenter, in your inbox

Data markets, alt data, and the AI training-data economy. No spam, unsubscribe anytime.

Discussion lives in the inline notes attached to article passages.