Open Source AI
(65 articles)AGI Ranker Audit: Every LLM Score Dropped 6-15 Points
A self-audit of the AGI Ranker leaderboard exposed scoring bias that inflated every model by 6-15 points. Here's what the correction actually revealed.
Apple SpeechAnalyzer vs Whisper: Benchmark Verdict
Apple's new SpeechAnalyzer API landed in iOS 26 with big claims. Benchmark data from Inscribe puts it head-to-head with Whisper and the old SFSpeechRecognizer....
Train a Kick Drum AI Model on 6GB VRAM: Full Linux Guide
A dusty GTX 1660 and a weekend are all you need. This tutorial walks through training a working kick drum diffusion model on 6GB of VRAM, from dataset prep to...
Qwen 3.7 Plus vs 3.6 Plus: 7 Real Upgrades in 2026
A no-fluff breakdown of what actually changed between Qwen 3.7 Plus and Qwen 3.6 Plus, from reasoning gains to pricing shifts and coding wins.
DeepSeek V4 Pro vs V3: 7 Upgrades That Matter
DeepSeek V4 Pro replaces V3 with 1M-token context, a 1.6T-parameter MoE, and native reasoning modes. Here's which upgrades matter — and where V3 still wins on...
Talos-XII: Hand-Written Rust Autograd Hits 10k Sims/Sec
A solo-built Rust autograd stack with custom SIMD dispatch models gacha probabilities at 10k+ sims per second. Here's what the benchmarks reveal about...
7 Open-Source Claude Desktop Alternatives Worth Trying
Rowboat, LibreChat, Open WebUI, Jan, and more: seven serious open-source Claude Desktop alternatives ranked for 2026, with honest takes on each.
DeepSeek V4-Flash vs V3.2: 7 Real Differences That Matter
A hands-on look at DeepSeek V4-Flash vs V3.2. What actually changed in speed, coding, context, and pricing, and whether the upgrade is worth it for your...
ScarfBench: IBM's Brutal Test for Java Migration AI
IBM Research's ScarfBench puts AI coding agents through real enterprise Java framework migrations. The results show a big gap between demo-day hype and...
Senior SWE-Bench: The Benchmark That Humbles AI Agents
Snorkel AI's new Senior SWE-Bench flips the script on coding agents by testing them as senior engineers. Top models land in the 40-55% range on basic solves...
FFASR Leaderboard: ASR Benchmarked on Real-World Audio
Treble Technologies and Hugging Face just dropped the FFASR Leaderboard, a far-field ASR benchmark that exposes how badly clean-audio scores have been lying to...
DeepSWE Benchmark: 91 Repos, 5 Languages, Zero Leaks
DeepSWE is a fresh contamination-free coding benchmark spanning 91 repos and 5 languages. Here's what the numbers say about frontier coding agents.