GSK Pays Up to $110M for Relation’s Lab-Grown AI Training Data

GSK signed a deal worth up to $110 million with Relation Therapeutics on July 30, 2026, Reuters reported, to fund automated-lab generation of multi-omic cell perturbation data for Relation's MORGAN…

Is this a drug-discovery deal or a data-procurement contract wearing a pharma partnership’s clothes? The evidence says the latter. GSK’s agreement with Relation Therapeutics, worth up to $110 million and reported by Reuters on July 30, 2026, funds Relation’s automated laboratories to churn out what CEO David Roblin calls “petascale, high-resolution multi-omic perturbation datasets,” according to BioXconomy — the raw material for MORGAN, Relation’s new foundation model of cellular response. GSK isn’t buying a molecule; it’s buying throughput on a data factory it doesn’t own, betting that superhuman-consistency lab automation beats the messy, inconsistent perturbation data that’s publicly available or generated in-house.

The deal fits a pattern GlobalData flagged to Pharmaceutical Technology: total AI-partnership value in pharma rose 120% year-on-year between 2024 and 2025, and per GEN Engineering News, 2026 is already stacking up as pharma’s “infrastructure moment,” with Lilly backing Chai Discovery and a $1 billion Nvidia lab, Pfizer backing Boltz, and Novo Nordisk wiring OpenAI into R&D end to end. Relation itself is stacking buyers with distinct jobs — GSK for data and model-building, Novartis for target discovery, Deerfield for company creation — which is itself a tell about how sellers are learning to price the same underlying asset differently depending on what the buyer actually wants: raw data, validated targets, or spun-out ventures.

When a $110 million pharma check buys lab automation instead of a lead compound, the market has quietly repriced experimental data above the models trained on it.

Watch whether GSK’s payment structure ties tranches to data volume or dataset quality rather than clinical milestones — that would confirm pharma is now underwriting data infrastructure the way tech buyers underwrite GPU clusters, and would tell rival biotechs exactly what a petabyte of consistent perturbation data is worth on the open market.

We're generating petascale, high-resolution multi-omic perturbation datasets through automated laboratories designed specifically for high-throughput cellular experiments with superhuman consistency. That enables us to produce data at a quality, scale and reproducibility that conventional laboratory approaches simply can't achieve.

Reuters

Read the full story at Reuters →

The Data Commenter, in your inbox

Data markets, alt data, and the AI training-data economy. No spam, unsubscribe anytime.

Discussion lives in the inline notes attached to article passages.