Why would Google pay north of $1.5 billion for a startup that raised just $9.1 million at a $500 million valuation four months earlier? Because this deal, as reported by businessinsider.com on August 5, 2026, isn’t really pricing Mechanize the company — it’s pricing the scarcity of engineers who can build evaluation environments for coding agents, and Google is buying its way around that scarcity the same way it has before.
The structure is the tell. Google isn’t acquiring Mechanize outright; it’s hiring a slice of the talent and taking a non-exclusive license, the exact hybrid playbook it ran on Windsurf last year for $2.4 billion in licensing fees, according to pymnts.com. That precedent is instructive: both deals let Google grab capability while sidestepping the antitrust review a straight acquisition would invite, as Business Insider Africa and finance.biggo.com both note. A roughly 3x markup on Mechanize’s April 2025 valuation, if the $1.5 billion-plus figure holds, looks less like a valuation call and more like talent-scarcity pricing in a coding-agent arms race where OpenAI’s Codex and Anthropic’s Claude Code have already won developer mindshare.
Google isn’t buying Mechanize’s product — it’s buying out the queue for people who can grade AI at its own weakest skill.
The talks land awkwardly alongside Jeff Dean’s departure after 27 years to co-found Discovery Loop, taking fellow veterans Oriol Vinyals, Quoc Le, and Sanjay Ghemawat with him — a defection Stocktwits reports helped drive Alphabet’s stock down nearly 4% on the news, its worst single-day drop in two weeks. Add Noam Shazeer’s second exit to OpenAI and DeepMind’s John Jumper leaving for Anthropic, and Google’s pattern is clear: it is bleeding senior research talent even as it opens the checkbook for outside teams to plug coding gaps. Wall Street isn’t panicking yet — GOOGL still carries a Strong Buy consensus and a price target implying 17.4% upside, per biggo.com — but the market for evaluation and benchmarking talent is now openly priced in the billions, and every AI lab with a checkbook is watching how this one closes.
We will achieve this by creating simulated environments and evaluations that capture the full scope of what people do at their jobs. This includes using a computer, completing long-horizon tasks that lack clear criteria for success, coordinating with others, and reprioritizing in the face of obstacles and interruptions.