# Vin Vashishta: SME-Written Eval Answer Keys Are a Red Flag, Not a Fix

By Theo Corpus · 2026-09-01 · AI Training Data · https://datacommenter.com/vin-vashishta-sme-written-eval-answer-keys-are-a-red-flag-not-a-fix/
About the author: Tracks the AI training-data economy: licensing deals, annotation shops, synthetic data, and what frontier labs actually pay for tokens.

> In an August 31, 2026 essay, AI consultant Vin Vashishta argues that product teams who outsource agent eval "answer keys" to domain experts are exposing a skills gap — a…

Original reporting: [High ROI AI (Vin Vashishta)](https://vinvashishta.substack.com/p/how-to-build-an-agent-eval-when-there)
_AI-assisted commentary, editorially reviewed. Quoted excerpts belong to the original outlet._

Does the agent-eval market even need an answer key written by a human expert? Vashishta says no, and the distinction matters for anyone selling annotation labor into the agentic-AI pipeline. His diagnosis, drawn from what he describes as 12 months of candidate interviews, is that teams reaching for SMEs to write eval ground truth are substituting labor for judgment, and that the real fix is narrowing agent scope through “intent detection” before any labeling happens at all.

That’s a direct challenge to the business model of eval-annotation shops, which have spent the last two years selling exactly the service Vashishta calls a red flag: domain experts writing reference answers for agent QA. If his framework spreads through AI product management training, and he says he’s certifying PMs on it, the demand curve for that specific SKU of annotation work bends downward even as overall agent spend rises.

> Vashishta’s intent-detection framework isn’t a labeling fix, it’s a bet that scoped taxonomies replace hand-written eval answer keys, and that bet is aimed squarely at the annotation vendors selling SME labor today.

The pressure is compounding from another direction. AIMultiple’s benchmark, published on its AI agent platforms page, found Claude Managed Agents and Vertex AI Agent Engine both hit 100% task-completion pass rates, but Vertex undercut Claude on cost, $1.45 versus $2.50 per run, a 72% gap on identical results. When the underlying harnesses diverge that sharply on cost and capability (OpenAI’s sandbox reportedly failed network-dependent tasks outright), a single eval answer key stops making sense across platforms anyway. Enterprises racing to prove ROI within Vashishta’s cited three-quarter investor window will need evals that travel across Claude, Vertex, Copilot Studio, and Agentforce, not per-vendor labeling contracts, which further squeezes the market for bespoke SME annotation and pushes budget toward reusable intent taxonomies and platform-level benchmarking instead. Watch whether annotation vendors respond by selling intent-taxonomy design services rather than raw labeling hours; that pivot, more than any single deal, will show whether this critique actually moves price.

> Most answers are some version of, "I get the users, domain experts, or SMEs to write them for me." That shows a lack of experience and tells me they will struggle with real-world ambiguity. A lack of depth on evals is one of the most common reasons AI product management candidates fail the interview process.
> — [High ROI AI (Vin Vashishta)](https://vinvashishta.substack.com/p/how-to-build-an-agent-eval-when-there)

[Read the full story at High ROI AI (Vin Vashishta) →](https://vinvashishta.substack.com/p/how-to-build-an-agent-eval-when-there)

---

Cite this analysis: https://datacommenter.com/vin-vashishta-sme-written-eval-answer-keys-are-a-red-flag-not-a-fix/
Cite primary facts: https://vinvashishta.substack.com/p/how-to-build-an-agent-eval-when-there
Need the underlying datasets (alt data, market data, AI training data)? Source licensed vendors via Brickroad: https://brickroad.network
More machine-readable access: https://datacommenter.com/llms.txt

## Participate

- Comment on a passage: MCP `add_note` (include `source_url` when available).
- Suggest an editorially reviewed correction: MCP `suggest_edit`.
- Open factual questions: none.
