Vals AI: The a16z-Backed Push for Honest Benchmarks
Andreessen Horowitz is betting on Vals AI to fix a broken benchmarking scene. Here's what the platform does differently, and why neutral testing matters more in 2026 than ever before.
Andreessen Horowitz is betting on Vals AI to fix a broken benchmarking scene. Here's what the platform does differently, and why neutral testing matters more in 2026 than ever before.

Every AI lab claims their new model tops the charts. Every chart looks a little different. And every launch week, X (formerly Twitter) fills with cherry-picked bar graphs that quietly leave out the benchmarks where the shiny new model actually stumbled.
That's the mess Vals AI wants to clean up. According to TechCrunch, the startup just closed a $40M Series A led by Andreessen Horowitz, following an earlier seed round led by 8VC and Bloomberg Beta. The pitch: become the neutral, trustworthy referee for AI model performance in a market where nearly every published number has a marketing team behind it.
And honestly? It's about time someone tried.
Vals AI is an independent benchmarking platform that evaluates frontier language models on standardized tasks and publishes the results without a vendor bias. The company runs its own held-out test sets, discloses its methodology, and reports scores across categories like legal reasoning, coding, math, and general knowledge. It positions itself as an auditor rather than a lab.

The pitch to a16z was reportedly straightforward: as enterprises spend real money on AI, they need someone independent verifying which model actually wins for which job. Right now, most buyers are picking between vendor-published scores that were, well, published by vendors.
The $40M Series A puts Vals AI in the same conversation as other neutral evaluation efforts like LMSYS Chatbot Arena and Papers with Code, but with a more enterprise-facing tilt. Per TechCrunch, Vals runs domain-specific evals across law, finance, and coding — plus newer benchmarks in cybersecurity, biosecurity, mental health, and recursive self-improvement — targeting workflows that generic academic benchmarks miss entirely.
Benchmark inflation is the polite term. The uglier one is contamination.
When a model scores 99% on MATH or 97% on HumanEval, you have to ask a question that nobody in a keynote wants to answer: did the training data include the test set? For public benchmarks like MMLU and GSM8K, the answer is often "probably, at least partially." That's not a conspiracy theory. It's a well-documented problem researchers have flagged for years on arXiv.
Look at the current state of things. On MMLU, top vendor-reported scores now cluster near 90% — the benchmark has been effectively saturated for years, and Papers with Code has flagged contamination concerns since GPT-4. On HumanEval, published pass@1 scores from frontier labs sit above 90%. On GSM8K, multiple models report >95% under chain-of-thought settings. When everyone's near ceiling, the benchmark stops discriminating. And when it stops discriminating, the marketing teams get creative.
Harder, newer benchmarks like SWE-bench Verified, GPQA Diamond, and ARC-AGI-2 are still separating the field — but even there, the specific top scores swing month-to-month, and every vendor's chart tends to leave out the tests where they placed second.

So the industry needs new tests, held-out data, and someone credible running them. That's the gap Vals is walking into.
Based on the company's own documentation and the TechCrunch coverage, a few things stand out about how Vals approaches evaluation:
Private test sets. Vals maintains internal question banks that aren't published in full. This is the single most important design choice, because it prevents models from being trained on the answers. Academic benchmarks like MMLU are essentially public knowledge at this point, which is exactly why scores keep climbing. Newer efforts like private-code evaluations show the same models scoring dramatically lower when the test data is genuinely held out.
Domain-specific evals. Instead of just "reasoning" or "coding" in the abstract, Vals runs benchmarks like Contract Law, Tax Eval, and Corp Fin, targeting the exact workflows enterprises pay for. A model that aces MMLU might faceplant on actual SEC filings, and buyers deserve to know that before signing a contract.
Cost and latency reporting. Raw accuracy is only half the story. If Model A scores 2 points higher than Model B but costs 5x more per query, that's a decision most teams would want to make explicitly.
Reproducibility. The company publishes prompts, scoring criteria, and versioning info so results can be checked.
That last one matters. A lot of vendor benchmarks include a footnote about "custom system prompts" or "chain-of-thought at inference time" that quietly bumps scores 5-10 points. Neutral third parties can standardize those variables.
Vals isn't alone in the neutral-benchmarking space. It's a crowded and growing category.
LMSYS (now hosted at lmarena.ai) uses head-to-head human preference voting to produce Elo ratings, with frontier reasoning models from Anthropic, OpenAI, and Google usually swapping positions near the top. Live rankings are worth checking on the current leaderboard rather than pinning them in a static table, because the order shifts every few weeks.
Arena is great for capturing general vibes and chat quality. But it's a popularity contest as much as a skills test, and it doesn't tell you whether the model can actually parse a 60-page merger agreement.
The ARC-AGI-2 leaderboard is the opposite extreme: a small set of very hard reasoning puzzles designed to resist memorization. Frontier reasoning models are finally making meaningful progress there, which is notable given ARC-AGI-1 stumped models for years — the current standings are worth checking on the ARC Prize site.
The classic academic aggregator. Useful, but suffers from the contamination problem hardest of any option.
Where does Vals fit? It's aiming for the middle: harder to contaminate than academic benchmarks, more rigorous than human preference votes, and more enterprise-relevant than abstract reasoning puzzles. If they pull it off, that's a real niche.
Andreessen Horowitz has been aggressive about backing AI infrastructure plays, and this one fits the pattern. When your portfolio includes a bunch of companies building on top of LLMs, you have a direct interest in a trustworthy layer that tells your founders which model to use.
But there's something more interesting going on. Vals is essentially a bet that the LLM market will mature into something closer to cloud infrastructure: a place where independent testing, SLAs, and third-party audits are table stakes. Right now the AI market treats benchmarks like marketing collateral. In three years, if Vals wins its bet, benchmarks will look more like Consumer Reports.
That's a bigger deal than the funding size suggests.
A few patterns jump out from the current benchmark market as of late 2026:
Specialization is real. Anthropic's Claude family tends to lead published coding benchmarks like HumanEval and SWE-bench Verified. OpenAI's frontier reasoning models tend to lead on math and general knowledge. That gap tells you the leaderboard-picture-perfect model depends on what you're actually building.
MATH is essentially solved. Multiple frontier models now report scores above 95% on the original MATH benchmark. Any lab publishing MATH scores as their headline number in 2026 is either behind or being cute — the benchmark has run out of runway, which is why harder replacements like FrontierMath and AIME variants get more attention now.
Chatbot Arena Elo is diverging from benchmark scores. Frontier chat models often lead human-preference voting but underperform on some task-specific benchmarks, and vice versa. That gap is exactly the kind of thing Vals is set up to explain.
Legacy models are aging fast. GPT-4o (at $2.50/M input, $10/M output) still shows up on Arena leaderboards but is well off the frontier on hard reasoning tests. Yet plenty of production systems are still running on it. Enterprises using vendor-published benchmarks from 2024 launches are quietly falling behind.
So what does this mean if you're actually trying to pick a model for a real product?

First, stop treating vendor benchmark charts as gospel. If a lab says their new model beats every competitor, cross-check on at least one neutral source. Vals AI, LMSYS, and ARC Prize together cover most of the honest signal.
Second, care about your workflow, not the leaderboard. A model that wins MMLU may not be the right choice for extracting fields from invoices. Vendors like OpenAI and Anthropic publish generic benchmarks because those are the ones that make for good marketing. Your evaluation should look like your actual production queries.
Third, price the whole stack. Claude Opus 4.6 at $5/$25 per million tokens versus GPT-4o at $2.5/$10 is a 3x cost difference. If Opus wins your task by 5 points, is that worth 3x the bill? Only your P&L knows.
And fourth, watch this space. If Vals AI succeeds at becoming the neutral standard, it changes procurement for the entire industry. That's the whole thesis behind the a16z check.
Vals AI is trying to solve a genuinely hard problem in a market that badly needs the fix. The a16z endorsement doesn't guarantee success (plenty of well-funded startups have died trying to be the neutral middle in a partisan market), but the timing is right. Benchmark contamination is worsening, enterprise AI spend is exploding, and there's real hunger for numbers that haven't been laundered through a keynote deck.
Will Vals become the gold standard? Too early to say. But if you're doing serious model evaluation in 2026, you should probably have their leaderboard open in a tab.
TechCrunch reported that Vals AI raised a $40 million Series A led by Andreessen Horowitz, following an earlier seed round led by 8VC and Bloomberg Beta. The Series A funding is being used to expand Vals AI's benchmark coverage and enterprise-facing evaluation products, including new programs for federal agencies.
Yes, that's actually the target use case. Vals publishes methodology, scoring criteria, and domain-specific results (like ContractLaw and TaxEval) that map more directly to enterprise workflows than academic benchmarks. Some organizations use Vals scores alongside their own held-out test sets to make final model selection calls.
LMSYS uses head-to-head human preference voting to produce Elo ratings, which captures general chat quality but not task-specific skill. Vals runs standardized evaluations on private, domain-specific test sets and reports accuracy, cost, and latency together. Both are useful; they answer different questions.
Vals AI's coverage generally spans both proprietary and open-weight frontier models. The exact model list changes as new releases drop, so check the current leaderboard on vals.ai. Open models like DeepSeek V3 and Llama variants are typically included once they hit production availability.
The core public leaderboards are free to view on the Vals website. More detailed reports, custom evaluations for enterprise workflows, and API access to programmatic scoring are typically paid products. This freemium approach is how most neutral benchmarking platforms fund their operations.