Benchmarks
AI model benchmark results and analysis — 35 articles

Real-SWE Benchmark: AI Models Struggle on Private Code
Real-SWE tests AI models on private enterprise codebases instead of public GitHub repos. The top frontier model scores...

ASR Benchmarks Are Broken: The Optimization Problem
Speech recognition leaderboards keep hitting record-low WER scores, but Hugging Face's new analysis shows how much of...

Pocket LLM Benchmarks: What Actually Runs on Your Phone
The Artificial Analysis mobile inference data shows a widening gap between what phones can theoretically run and what...

LoRA Speedrun: Wall-Clock Leaderboard Shaking Up Fine-Tuning
A new public leaderboard called LoRA Speedrun ranks fine-tuning techniques by raw wall-clock time. The results expose...

ASR Benchmark Gaming: How to Spot Overfitting in 2026
Hugging Face's new methodology shows how speech recognition models game leaderboards. The gap between benchmark WER and...

Production RAG on Open Models: The Numbers That Matter
A benchmark-driven look at production RAG with open models, hybrid retrieval, reranking, and RAGAS scoring. What...

AI Benchmark Saturation: Why the Scoreboards Broke in 2026
MMLU is capped at 93%. HumanEval is basically solved. A look at the data behind AI benchmark saturation and what...

Homebench: The Local LLM Benchmark Tool Worth Your Time
Homebench measures speed, memory, and quality for local LLMs on your own hardware. Here's what the numbers actually...

AGI Ranker Audit: Every LLM Score Dropped 6-15 Points
A self-audit of the AGI Ranker leaderboard exposed scoring bias that inflated every model by 6-15 points. Here's what...

LLM Agents Flop at Coordination: Inside the ALEM Benchmark
A new open-ended coordination benchmark tests 13 LLMs across communication, trading, crafting, and combat. Most agents...

Apple SpeechAnalyzer vs Whisper: Benchmark Verdict
Apple's new SpeechAnalyzer API landed in iOS 26 with big claims. Benchmark data from Inscribe puts it head-to-head with...

Talos-XII: Hand-Written Rust Autograd Hits 10k Sims/Sec
A solo-built Rust autograd stack with custom SIMD dispatch models gacha probabilities at 10k+ sims per second. Here's...