The headline number here isn’t about data licensing directly, but it’s about where enterprise dollars for context infrastructure are actually flowing — and the answer is bad news for standalone vector database vendors. VentureBeat’s survey of 101 enterprises, published July 16, 2026, shows OpenAI’s file search (40%) and Google’s Vertex AI Search (38%) have overtaken every purpose-built vector database in production use, with specialists like Elasticsearch/OpenSearch (20%) and pgvector (12%) trailing and category-defining names like Weaviate, Pinecone, Qdrant, and Milvus stuck in single digits.
For the data economy, the more consequential figure is the 57% of enterprises that traced a confident-but-wrong agent answer to thin or inconsistent business context in the last six months — with more than half of those saying it happened repeatedly. Retrieval is already the primary context source for 38% of organizations, nearly double the 21% relying on a governed semantic layer, which means the failure surface is concentrated exactly where enterprises are least confident. That gap is the real market signal: 58% say they’re building or running a governed semantic layer, but most admit it isn’t in production yet — a backlog of demand for curation, ontology-building, and context-governance work that annotation shops and data-infrastructure vendors should be racing to fill.
Enterprises are buying provider-native retrieval while budgeting, in the same breath, to build the independent context layer they say they still want.
The contradiction in buyer intent is worth watching closely. A plurality (36%) say they intend to keep best-of-breed standalone tools rather than consolidate onto a provider’s native stack, yet 57% plan to switch or add a provider within the year, and 34% expect hybrid retrieval to dominate by the end of 2026. That tension — stated preference for independence against actual usage tilting toward OpenAI and Google — is exactly the kind of gap that pricing power gets built on. Whoever controls the semantic layer that governs what gets retrieved, not just the retrieval pipe itself, is the one who eventually sets the terms for enterprise context spend; expect the next wave of this Pulse Research series to show whether that layer gets built by hyperscalers, RAG specialists, or a new class of context-governance vendors.
A majority of enterprises (57%) have already had an AI agent produce a confident, wrong answer they traced to bad context — wrong metrics, stale definitions, or missing documents — and more than half of those have seen it happen more than once.