Qdrant, the Berlin-and-New York vector database vendor, just gave away the one thing every buyer in this market claims to want and almost no single vendor can credibly produce alone: a 10-billion-document ground-truth benchmark. Per BigDATAwire’s September 3, 2026 report, Fineweb-10B pairs that corpus with roughly 120,000 ground-truth queries, according to Open Source For You, computed via more than a quadrillion brute-force distance calculations — a compute bill Qdrant is absorbing so the rest of the industry doesn’t have to. Constellation Research analyst Michael Ni told TechTarget the move is “an ecosystem and category-building move” that lets Qdrant subsidize evaluation costs while making itself “central to that evaluation.” That’s the real transaction here: Qdrant pays in GPU cycles and engineering time, and in return it gets to define the exam.
Qdrant isn’t just publishing a dataset — it’s trying to write the exam every rival vector database will eventually have to sit.
Who benefits beyond Qdrant matters too. Hugging Face, the Common Crawl Foundation, Vultr, and Alibaba get their names attached to a benchmark McKnight Consulting’s William McKnight says solves a genuine “massive compute barrier” problem, per TechTarget, lending it a credibility a single vendor’s internal test never earns. Enterprises evaluating Milvus, Weaviate, Redis, pgvector, or Postgres-based alternatives get a free, non-synthetic, production-scale test rather than building one in-house. But free isn’t neutral, and McKnight himself flagged the pattern: vendors tend to open-source benchmarks “they are likely good at competitively.”
A Benchmark Market Already Crowded and Contested
That skepticism is earned. PR Newswire carried a July 2026 release in which EnterpriseDB, using McKnight-commissioned benchmarks of its own, claimed EDB Postgres AI beats vector databases outright on speed and recall — a competing yardstick from a competing interest, built by the same firm now praising Qdrant’s approach. Meanwhile AIMultiple’s independent open-source benchmark found Qdrant trailing Redis on raw single-thread throughput (377 QPS versus 764) even as it led on other dimensions, a reminder that whoever supplies the dataset doesn’t automatically top the leaderboard. Watch whether Milvus, Weaviate, and Redis adopt Fineweb-10B as a shared standard or respond with their own datasets, and whether third parties running Qdrant’s own Supernova framework reproduce Qdrant’s favorable numbers or complicate them.
Vector database company Qdrant has released a dataset with 10 billion documents to help put large-scale vector search systems through their paces. Benchmarking these systems isn't easy, particularly as they get bigger. You need a large enough dataset to make the test meaningful, but you also need to know what the correct results should be.