Skip to content
S

Shadman Ahmed

Software Architect

Software architect and AI tools enthusiast. I test, benchmark, and review AI models and developer tools so you don't have to.

189

Articles

103,568

Total Views

332K

Words Written

All Articles (189 total)

DeepSeek V4 Pro vs V3: 7 Upgrades That Matter

DeepSeek V4 Pro replaces V3 with 1M-token context, a 1.6T-parameter MoE, and native reasoning modes. Here's which upgrades matter — and where V3 still wins on cost.

July 11, 2026 9 min 194comparisons

Talos-XII: Hand-Written Rust Autograd Hits 10k Sims/Sec

A solo-built Rust autograd stack with custom SIMD dispatch models gacha probabilities at 10k+ sims per second. Here's what the benchmarks reveal about tiny-model performance without PyTorch.

July 10, 2026 7 min 180benchmarks

7 Open-Source Claude Desktop Alternatives Worth Trying

Rowboat, LibreChat, Open WebUI, Jan, and more: seven serious open-source Claude Desktop alternatives ranked for 2026, with honest takes on each.

July 8, 2026 8 min 202listicles

Qwen 3.7 Max Review: The Best Coding Value of 2026?

An honest look at Qwen 3.7 Max for coding: benchmarks, pricing versus Claude and GPT, real-world agent workflows, and whether the Alibaba frontier model is finally worth switching to.

July 7, 2026 9 min 274reviews

DeepSeek V4-Flash vs V3.2: 7 Real Differences That Matter

A hands-on look at DeepSeek V4-Flash vs V3.2. What actually changed in speed, coding, context, and pricing, and whether the upgrade is worth it for your workflow.

July 6, 2026 8 min 253comparisons

REAP Explained: Real Coding Benchmarks From Live Agent Traffic

REAP mines production coding agent sessions to build execution-based benchmarks. On the Harvest benchmark it produced, frontier models solve 42.9%-58.2% — well below SWE-bench Verified scores.

July 5, 2026 7 min 211benchmarks

ScarfBench: IBM's Brutal Test for Java Migration AI

IBM Research's ScarfBench puts AI coding agents through real enterprise Java framework migrations. The results show a big gap between demo-day hype and production reality.

July 3, 2026 7 min 194benchmarks

Senior SWE-Bench: The Benchmark That Humbles AI Agents

Snorkel AI's new Senior SWE-Bench flips the script on coding agents by testing them as senior engineers. Top models land in the 40-55% range on basic solves and fall into the teens or twenties once quality-graded.

July 2, 2026 8 min 229benchmarks

Grok 4.3 vs Grok 4.20: 5 Real Differences That Matter

xAI shipped Grok 4.20 alongside Grok 4.3 with a rebuilt reasoning stack and agentic tool loop. Same 1M context, same base price — where does the switch actually pay off?

July 1, 2026 8 min 312comparisons

Grok 4.3 Review: Is xAI's Reasoning Worth $30/Month?

An honest look at Grok 4.3's Think mode, real-time X data, and reasoning benchmarks. Where it actually beats Claude and GPT-5.5, and where it doesn't.

June 28, 2026 9 min 341reviews

FFASR Leaderboard: ASR Benchmarked on Real-World Audio

Treble Technologies and Hugging Face just dropped the FFASR Leaderboard, a far-field ASR benchmark that exposes how badly clean-audio scores have been lying to us.

June 27, 2026 8 min 259benchmarks

DeepSWE Benchmark: 91 Repos, 5 Languages, Zero Leaks

DeepSWE is a fresh contamination-free coding benchmark spanning 91 repos and 5 languages. Here's what the numbers say about frontier coding agents.

June 25, 2026 8 min 303benchmarks
PreviousPage 5 of 16Next