Skip to content

Benchmarks

AI model benchmark results and analysis26 articles

Two people looking at data on a laptop screen

LLM Agents Flop at Coordination: Inside the ALEM Benchmark

A new open-ended coordination benchmark tests 13 LLMs across communication, trading, crafting, and combat. Most agents...

July 19, 20268 min
Apple SpeechAnalyzer vs Whisper: Benchmark Verdict

Apple SpeechAnalyzer vs Whisper: Benchmark Verdict

Apple's new SpeechAnalyzer API landed in iOS 26 with big claims. Benchmark data from Inscribe puts it head-to-head with...

July 18, 20267 min
Talos-XII: Hand-Written Rust Autograd Hits 10k Sims/Sec

Talos-XII: Hand-Written Rust Autograd Hits 10k Sims/Sec

A solo-built Rust autograd stack with custom SIMD dispatch models gacha probabilities at 10k+ sims per second. Here's...

July 10, 20267 min
REAP Explained: Real Coding Benchmarks From Live Agent Traffic

REAP Explained: Real Coding Benchmarks From Live Agent Traffic

REAP mines production coding agent sessions to build execution-based benchmarks. On the Harvest benchmark it produced,...

July 5, 20267 min
ScarfBench: IBM's Brutal Test for Java Migration AI

ScarfBench: IBM's Brutal Test for Java Migration AI

IBM Research's ScarfBench puts AI coding agents through real enterprise Java framework migrations. The results show a...

July 3, 20267 min
Senior SWE-Bench: The Benchmark That Humbles AI Agents

Senior SWE-Bench: The Benchmark That Humbles AI Agents

Snorkel AI's new Senior SWE-Bench flips the script on coding agents by testing them as senior engineers. Top models...

July 2, 20268 min
FFASR Leaderboard: ASR Benchmarked on Real-World Audio

FFASR Leaderboard: ASR Benchmarked on Real-World Audio

Treble Technologies and Hugging Face just dropped the FFASR Leaderboard, a far-field ASR benchmark that exposes how...

June 27, 20268 min
DeepSWE Benchmark: 91 Repos, 5 Languages, Zero Leaks

DeepSWE Benchmark: 91 Repos, 5 Languages, Zero Leaks

DeepSWE is a fresh contamination-free coding benchmark spanning 91 repos and 5 languages. Here's what the numbers say...

June 25, 20268 min
Rio3.5 vs Qwen3.7: Why This Viral Benchmark Smells Off

Rio3.5 vs Qwen3.7: Why This Viral Benchmark Smells Off

A tweet claims Rio de Janeiro's city government built an LLM that beats Qwen3.7. No paper, no leaderboard, no weights....

June 23, 20267 min
Agentic LLM Benchmark: Open Models On Real Tooling

Agentic LLM Benchmark: Open Models On Real Tooling

Hugging Face's new agentic benchmark stress-tests open models against your actual toolset. The results expose a gap...

June 18, 20268 min
Bilingual Voice Agents Hit a Wall: ASR Code-Switch Benchmark

Bilingual Voice Agents Hit a Wall: ASR Code-Switch Benchmark

Frontier ASR models stumble when customers mix two languages in one sentence. A new ServiceNow-AI benchmark exposes how...

June 10, 20268 min
ITBench-AA: Top AI Models Flunk Enterprise IT Tasks

ITBench-AA: Top AI Models Flunk Enterprise IT Tasks

IBM and Artificial Analysis just dropped ITBench-AA, the first real test of AI agents on enterprise IT work. Every...

June 3, 20268 min