Company Knowledge Bench: Why RAG Still Fails at Work
Kapa.ai's new benchmark tests retrieval on actual messy enterprise data, and the results expose how badly most RAG systems break outside clean academic datasets.
Kapa.ai's new benchmark tests retrieval on actual messy enterprise data, and the results expose how badly most RAG systems break outside clean academic datasets.

Most RAG benchmarks are a lie. Not deliberately, but functionally. They test clean Wikipedia passages, synthetic Q&A pairs, or hand-curated technical docs that look nothing like the ticket graveyards, duplicated Confluence pages, and half-dead Slack threads real employees actually search through.
That's the gap Kapa.ai just tried to close with the Company Knowledge Bench, a retrieval benchmark built from genuinely messy real-world company knowledge. And the numbers it produces are, not gonna lie, a bit of a reality check for anyone shipping agents into the enterprise.
The Company Knowledge Bench is a retrieval benchmark for AI agents that evaluates how well retrieval systems surface the right chunks from a company's actual indexed corpus. It uses 1,000 eval cases built from real production queries across Kapa's deployed base, each paired with a snapshot of the corpus and a retrieval criterion that specifies which chunks are valid to return. The point is to measure performance where it hurts: duplicates, outdated pages, conflicting sources, and ambiguous queries.
That's a very different exercise from running MTEB or BEIR against polished academic datasets. Those are useful for component-level comparisons. They're not useful for predicting whether your support agent will cite the right runbook.
If you've shipped a RAG system in production, you already know the drill. You pick an embedding model with a strong BEIR score. You wire up a reranker. Evaluation on your clean test set looks great. Then real users ask real questions and the system confidently cites a 2023 doc that was deprecated six months ago.
The problem isn't the models. It's the evaluation. Public benchmarks almost always assume:
Real company knowledge violates every one of those assumptions. A single question about "rate limits" might have fifteen plausible sources, three of which contradict each other, and only one of which reflects the current API version. Kapa's write-up argues that public retrieval benchmarks are mostly too narrow (built around one domain like law or medicine) or too artificial (built from synthetic documents and questions), and that this is exactly why they built the benchmark in-house against real production data.
The benchmark's labelling handbook calls out specific situations that trip retrieval systems. Kapa doesn't publish per-category score drops, but the categories themselves describe what the benchmark is probing for:
| Failure mode | What it looks like |
|---|---|
| Duplicate / near-duplicate docs | Multiple versions of the same guide |
| Outdated but still indexed content | Deprecated API docs resurfacing |
| Conflicting sources | Support ticket contradicts official docs |
| Ambiguous queries | "How do I set this up?" with no context |
| Multi-source answers | Answer requires combining or choosing between docs |
The benchmark's definition of a correct retrieval is a minimal set of chunks that completely answers the query using the highest-authority, most-current sources. Returning a stale page or a lower-authority duplicate counts against you even if it technically contains the answer.
The methodology is more interesting than the usual "embed, retrieve, compute nDCG" loop. Each eval case is a real production query, a snapshot of the corpus as it was when the query was asked, and a boolean retrieval criterion specifying which chunks are valid to return.
Kapa built the dataset through:
Grading rewards completeness (did you return enough chunks to answer?), minimality (did you avoid returning extras?), and source preference (did you pick the authoritative, current source over duplicates and outdated versions?).
The benchmark itself is not public, since it's built from real production data. The methodology is described in enough detail that teams can reproduce it on their own corpus.
Kapa evaluated seven retrievers on the same ingested data, with 1,000 eval cases. Score is on a 0-1 scale where higher is better. These are the numbers Kapa reported:
| Retrieval approach | Retrieval score | Latency | Precision |
|---|---|---|---|
| Hybrid search | 0.41 | 0.4 s | 8% |
| Hybrid search + reranker | 0.50 | 0.7 s | 9% |
| Query decomposition + hybrid + rerank | 0.56 | 1.8 s | 10% |
| Kapa Default (optimized fixed pipeline) | 0.61 | 3.3 s | 12% |
| Agent + grep (small model, Luna) | 0.54 | 13 s | 4% |
| Agent + grep (large model, Sol) | 0.61 | 17 s | 4% |
| Kapa Deep (agentic retrieval) | 0.65 | 5.1 s | 26% |
Hybrid search uses gemini-embedding-001 with BM25. The reranker is Voyage AI's rerank-2. Agentic grep uses OpenAI's gpt-6.1-sol or gpt-6.1-luna with full-text and regex search as its only tool.
Two takeaways jump out, both of which Kapa calls out directly. First, adding a reranker is the cheapest big win in traditional RAG. Going from plain hybrid search to hybrid + reranker lifts the score from 0.41 to 0.50 for a third of a second more latency. Second, agentic retrieval wins on quality but at a steep cost. The grep agent with a strong model matches an optimized fixed pipeline on score (0.61) but takes five times as long. Kapa's own Deep mode (optimized agentic retriever) scores 0.65 in about five seconds, and the pipeline bias here is explicit — all seven retrievers were built by Kapa and run on their ingestion.
Agentic retrieval is one of those phrases that gets thrown around without much substance. In Kapa's benchmark it's concrete. The agent:
It's slower. It costs more tokens. The grep agents in the benchmark return around 40,000 tokens per query, roughly four times a fixed pipeline, and the large-model grep agent spends about $0.17 in orchestrator model calls per query. An optimized agentic retriever (Kapa Deep) sidesteps most of that, returning about 5,000 tokens per query at the highest precision in the benchmark.
A few findings pushed back against conventional RAG advice.
A reranker is the cheapest big win. Going from hybrid search alone to hybrid plus a reranker lifted the score from 0.41 to 0.50 with only 0.3 seconds of added latency. If you change one thing in a basic RAG pipeline, this is it.
Query decomposition helps, but at a latency and token price. Splitting the query first into sub-queries added 0.06 in score over hybrid + rerank, but more than doubled the latency to 1.8 seconds and added an orchestrator model call per query.
An agentic loop is not automatically better. The small-model grep agent (0.54) actually lost to query decomposition (0.56) in a seventh of the time. For agentic retrieval, the orchestrator model's strength matters a lot — swapping the large model for the smaller one cost 0.07 in score.
Precision varies more than score. Kapa Deep returned about 5,000 tokens per query at 26% precision (share of chunks actually needed). The grep agents returned ~40,000 tokens at just 4% precision. Even when end score is similar, token efficiency affects what the downstream model pays to read the result.
The honest implication is that if you've evaluated your RAG system only on a clean internal eval set, you probably have no idea how it performs. Enterprise retrieval evaluation needs to look like enterprise data, which means messy, redundant, and partially contradictory. If you want numbers on how production RAG on open models actually behaves at scale, that's a different piece of the same puzzle.
A few concrete moves worth considering:
And for anyone building retrieval infrastructure, the Company Knowledge Bench is a useful target. According to the Kapa.ai team, the dataset itself isn't public since it's built from real production data, but the methodology (eval cases as query + corpus + boolean retrieval criterion, graded on completeness, minimality, and source preference) is reproducible on your own corpus. That's probably the more useful move anyway, since the whole point is that generic benchmarks don't predict your performance.
Retrieval is having its "evaluation gap" moment, the same one that hit LLM evaluation two years ago when everyone realized MMLU scores didn't predict real-world helpfulness. Benchmarks like SWE-bench Verified solved that for coding agents by using real GitHub issues. The Company Knowledge Bench is trying to do the same for enterprise retrieval.
If it catches on, expect a flurry of "we top the Company Knowledge Bench" marketing from retrieval vendors over the next six months. We've already seen this pattern with framework-level comparisons like LangChain vs LlamaIndex vs Haystack, where vendor-run numbers rarely survive contact with your corpus. Take those claims with the same skepticism you'd apply to any self-reported benchmark — including Kapa's, where all seven retrievers in the published results were built by Kapa and run on their ingestion. But the underlying shift (toward real, messy, enterprise-grade retrieval evaluation) is overdue and genuinely useful.
The best retrieval system is the one that works on your actual data. Everything else is a leaderboard.
No. The benchmark is private because it's built from real production queries and customer corpora. Kapa has published the methodology in enough detail that teams can reproduce it on their own data (eval cases as query + corpus snapshot + a boolean retrieval criterion graded on completeness, minimality, and source preference). That's actually the recommended approach since the whole point is domain-specific evaluation.
BEIR and MTEB evaluate retrieval components on clean, mostly single-answer datasets like Natural Questions or FiQA. Company Knowledge Bench evaluates end-to-end retrieval pipelines on noisy, redundant corpora with multiple valid answers per query and source-preference rules. They're complementary: use BEIR to compare embedding models, use something like CKB to predict production performance.
In Kapa's reported numbers, their own Kapa Deep agentic retriever scored highest at 0.65 on a 0-1 scale, followed by Kapa Default and a large-model grep agent tied at 0.61. Plain hybrid search scored 0.41 and adding a reranker lifted it to 0.50. All seven retrievers were built by Kapa, so treat the ordering accordingly — but the direction is clear: adding a reranker is the cheapest single improvement, and agentic retrieval wins on quality at the cost of latency and tokens.
In Kapa's results, hybrid + reranker scored 0.50, while their optimized agentic retriever (Kapa Deep) scored 0.65 at about 5 seconds per query. For high-stakes support or compliance use cases the extra latency and cost of agentic retrieval are usually worth it. For low-stakes autocomplete, a reranker is likely enough. The small-model grep agent in the benchmark actually lost to simple query decomposition, so an agentic loop is not automatically better.
It depends heavily on the orchestrator model and how many search iterations it runs. In Kapa's benchmark, their large-model grep agent spent about $0.17 per query in orchestrator calls and returned around 40,000 tokens, versus about $0.01 for Kapa Deep (which returned ~5,000 tokens). If you plug in a premium model like Claude Opus 4.6 at $5/$25 per MTok as your orchestrator, costs climb proportionally. Caching, query deduplication, and returning fewer tokens all help.