Showing 35 benchmarks articles
BenchmarksREAP mines production coding agent sessions to build execution-based benchmarks. On the Harvest benchmark it produced,...
BenchmarksIBM Research's ScarfBench puts AI coding agents through real enterprise Java framework migrations. The results show a...
BenchmarksSnorkel AI's new Senior SWE-Bench flips the script on coding agents by testing them as senior engineers. Top models...
BenchmarksTreble Technologies and Hugging Face just dropped the FFASR Leaderboard, a far-field ASR benchmark that exposes how...
BenchmarksDeepSWE is a fresh contamination-free coding benchmark spanning 91 repos and 5 languages. Here's what the numbers say...
BenchmarksA tweet claims Rio de Janeiro's city government built an LLM that beats Qwen3.7. No paper, no leaderboard, no weights....
BenchmarksHugging Face's new agentic benchmark stress-tests open models against your actual toolset. The results expose a gap...
BenchmarksFrontier ASR models stumble when customers mix two languages in one sentence. A new ServiceNow-AI benchmark exposes how...
BenchmarksIBM and Artificial Analysis just dropped ITBench-AA, the first real test of AI agents on enterprise IT work. Every...
BenchmarksGoogle's Antigravity 2.0 just posted the strongest autonomous result on ModelRift's OpenSCAD LLM benchmark, beating...
BenchmarksClaude Opus 4.6 reaches 81.4% on SWE-bench Verified per Anthropic, but raw HumanEval scores tell a different story. A...
BenchmarksAggregated 2026 benchmark data across three RAG frameworks reveals a clear split: LangChain wins ecosystem, LlamaIndex...