Skip to content

Benchmarks

AI model benchmark results and analysis35 articles

Real-SWE Benchmark: AI Models Struggle on Private Code

Real-SWE Benchmark: AI Models Struggle on Private Code

Real-SWE tests AI models on private enterprise codebases instead of public GitHub repos. The top frontier model scores...

September 14, 20267 min
ASR Benchmarks Are Broken: The Optimization Problem

ASR Benchmarks Are Broken: The Optimization Problem

Speech recognition leaderboards keep hitting record-low WER scores, but Hugging Face's new analysis shows how much of...

September 4, 20267 min
Pocket LLM Benchmarks: What Actually Runs on Your Phone

Pocket LLM Benchmarks: What Actually Runs on Your Phone

The Artificial Analysis mobile inference data shows a widening gap between what phones can theoretically run and what...

August 31, 20267 min
LoRA Speedrun: The Wall-Clock Leaderboard Shaking Up Fine-Tuning

LoRA Speedrun: Wall-Clock Leaderboard Shaking Up Fine-Tuning

A new public leaderboard called LoRA Speedrun ranks fine-tuning techniques by raw wall-clock time. The results expose...

August 31, 20268 min
ASR Benchmark Gaming: How to Spot Overfitting in 2026

ASR Benchmark Gaming: How to Spot Overfitting in 2026

Hugging Face's new methodology shows how speech recognition models game leaderboards. The gap between benchmark WER and...

August 23, 20268 min
Production RAG on Open Models: The Numbers That Matter

Production RAG on Open Models: The Numbers That Matter

A benchmark-driven look at production RAG with open models, hybrid retrieval, reranking, and RAGAS scoring. What...

August 21, 20268 min
AI Benchmark Saturation: Why the Scoreboards Broke in 2026

AI Benchmark Saturation: Why the Scoreboards Broke in 2026

MMLU is capped at 93%. HumanEval is basically solved. A look at the data behind AI benchmark saturation and what...

August 5, 20267 min
Homebench: The Local LLM Benchmark Tool Worth Your Time

Homebench: The Local LLM Benchmark Tool Worth Your Time

Homebench measures speed, memory, and quality for local LLMs on your own hardware. Here's what the numbers actually...

August 4, 20267 min
AGI Ranker Audit: Every LLM Score Dropped 6-15 Points

AGI Ranker Audit: Every LLM Score Dropped 6-15 Points

A self-audit of the AGI Ranker leaderboard exposed scoring bias that inflated every model by 6-15 points. Here's what...

July 31, 20265 min
Two people looking at data on a laptop screen

LLM Agents Flop at Coordination: Inside the ALEM Benchmark

A new open-ended coordination benchmark tests 13 LLMs across communication, trading, crafting, and combat. Most agents...

July 19, 20268 min
Apple SpeechAnalyzer vs Whisper: Benchmark Verdict

Apple SpeechAnalyzer vs Whisper: Benchmark Verdict

Apple's new SpeechAnalyzer API landed in iOS 26 with big claims. Benchmark data from Inscribe puts it head-to-head with...

July 18, 20267 min
Talos-XII: Hand-Written Rust Autograd Hits 10k Sims/Sec

Talos-XII: Hand-Written Rust Autograd Hits 10k Sims/Sec

A solo-built Rust autograd stack with custom SIMD dispatch models gacha probabilities at 10k+ sims per second. Here's...

July 10, 20267 min