Showing 27 benchmarks articles
BenchmarksA self-audit of the AGI Ranker leaderboard exposed scoring bias that inflated every model by 6-15 points. Here's what...
BenchmarksA new open-ended coordination benchmark tests 13 LLMs across communication, trading, crafting, and combat. Most agents...
BenchmarksApple's new SpeechAnalyzer API landed in iOS 26 with big claims. Benchmark data from Inscribe puts it head-to-head with...
BenchmarksA solo-built Rust autograd stack with custom SIMD dispatch models gacha probabilities at 10k+ sims per second. Here's...
BenchmarksREAP mines production coding agent sessions to build execution-based benchmarks. On the Harvest benchmark it produced,...
BenchmarksIBM Research's ScarfBench puts AI coding agents through real enterprise Java framework migrations. The results show a...
BenchmarksSnorkel AI's new Senior SWE-Bench flips the script on coding agents by testing them as senior engineers. Top models...
BenchmarksTreble Technologies and Hugging Face just dropped the FFASR Leaderboard, a far-field ASR benchmark that exposes how...
BenchmarksDeepSWE is a fresh contamination-free coding benchmark spanning 91 repos and 5 languages. Here's what the numbers say...
BenchmarksA tweet claims Rio de Janeiro's city government built an LLM that beats Qwen3.7. No paper, no leaderboard, no weights....
BenchmarksHugging Face's new agentic benchmark stress-tests open models against your actual toolset. The results expose a gap...
BenchmarksFrontier ASR models stumble when customers mix two languages in one sentence. A new ServiceNow-AI benchmark exposes how...