Benchmarks
AI model benchmark results and analysis — 26 articles

LLM Agents Flop at Coordination: Inside the ALEM Benchmark
A new open-ended coordination benchmark tests 13 LLMs across communication, trading, crafting, and combat. Most agents...

Apple SpeechAnalyzer vs Whisper: Benchmark Verdict
Apple's new SpeechAnalyzer API landed in iOS 26 with big claims. Benchmark data from Inscribe puts it head-to-head with...

Talos-XII: Hand-Written Rust Autograd Hits 10k Sims/Sec
A solo-built Rust autograd stack with custom SIMD dispatch models gacha probabilities at 10k+ sims per second. Here's...

REAP Explained: Real Coding Benchmarks From Live Agent Traffic
REAP mines production coding agent sessions to build execution-based benchmarks. On the Harvest benchmark it produced,...

ScarfBench: IBM's Brutal Test for Java Migration AI
IBM Research's ScarfBench puts AI coding agents through real enterprise Java framework migrations. The results show a...

Senior SWE-Bench: The Benchmark That Humbles AI Agents
Snorkel AI's new Senior SWE-Bench flips the script on coding agents by testing them as senior engineers. Top models...

FFASR Leaderboard: ASR Benchmarked on Real-World Audio
Treble Technologies and Hugging Face just dropped the FFASR Leaderboard, a far-field ASR benchmark that exposes how...

DeepSWE Benchmark: 91 Repos, 5 Languages, Zero Leaks
DeepSWE is a fresh contamination-free coding benchmark spanning 91 repos and 5 languages. Here's what the numbers say...

Rio3.5 vs Qwen3.7: Why This Viral Benchmark Smells Off
A tweet claims Rio de Janeiro's city government built an LLM that beats Qwen3.7. No paper, no leaderboard, no weights....

Agentic LLM Benchmark: Open Models On Real Tooling
Hugging Face's new agentic benchmark stress-tests open models against your actual toolset. The results expose a gap...

Bilingual Voice Agents Hit a Wall: ASR Code-Switch Benchmark
Frontier ASR models stumble when customers mix two languages in one sentence. A new ServiceNow-AI benchmark exposes how...

ITBench-AA: Top AI Models Flunk Enterprise IT Tasks
IBM and Artificial Analysis just dropped ITBench-AA, the first real test of AI agents on enterprise IT work. Every...