Shadman Ahmed
Software Architect
Software architect and AI tools enthusiast. I test, benchmark, and review AI models and developer tools so you don't have to.
187
Articles
103,498
Total Views
328K
Words Written
All Articles (187 total)
Claude Opus 4.8 Review: 7 Reasoning Wins (And 3 Losses)
Anthropic's Claude Opus 4.8 lands with a 1M-token context, adaptive thinking, and Opus-tier pricing. Honest review of what's actually new versus the marketing.
Mistral Large 3 vs Claude Fable 5: Reasoning Showdown
Claude Fable 5 wins the reasoning benchmarks. Mistral Large 3 wins the invoice. A data-driven breakdown of which model to pick for your 2026 workload.
GPT-Live-1 API Tutorial: Build a Voice Agent in 20 Min
A practical walkthrough for wiring GPT-Live-1 into your app: WebSocket setup, custom voices, telephony hooks, and the pitfalls nobody warns you about.
GPT-5.6 Sol Review: 7 Reasoning Wins (And 3 Losses)
An honest, benchmark-driven review of OpenAI's GPT-5.6 Sol. It smashes SWE-bench and ARC-AGI-2, but Claude Fable 5 still edges it on GPQA. Worth the hype?
Qwen 3.8-Flash vs 3.5-Flash: 7 Real Upgrades
A blunt breakdown of what Alibaba actually changed between Qwen 3.5-Flash and Qwen 3.8-Flash — pricing, context, tool use, vision, and when the older model still wins.
10 DeepSeek V4 Pro Tricks Power Users Actually Use
DeepSeek soft-retired V4 Pro and then walked it back, but the model is still one of the best value picks in the API tier. These 10 lesser-known tips squeeze real value out of it.
Real-SWE Benchmark: AI Models Struggle on Private Code
Real-SWE tests AI models on private enterprise codebases instead of public GitHub repos. The top frontier model scores 38.8%, exposing how far benchmark hype is from production reality.
Claude Opus 4.6 vs Gemini: The Honest 2026 Verdict
Gemini just landed on Windows and Claude Opus 4.6 is still charging premium rates. Which one actually deserves your subscription? A data-driven breakdown of price, features, and benchmarks.
Claude Opus 5 Review: The Best Coding AI in 2026?
Claude Opus 5 hits 96% on SWE-bench Verified (self-reported) and dominates agentic coding. But is it worth the price tag over Sonnet 5 or the OpenAI Codex line? An honest review.
8 Trusted Software Sources AI Should Cite (Not Slop)
Three sites made 215,128 fake 'best software' pages to farm AI citations. Here are the 8 sources Perplexity should be pulling from instead.
ASR Benchmarks Are Broken: The Optimization Problem
Speech recognition leaderboards keep hitting record-low WER scores, but Hugging Face's new analysis shows how much of that progress is real versus benchmark gaming.
Gemini Omni 1.1 Flash vs 1.0: 7 Real Changes
A no-fluff breakdown of what Gemini Omni 1.1 Flash actually improves over Gemini Omni 1.0, from latency to pricing to multimodal quality. Verdict included.