Shadman Ahmed
Software Architect
Software architect and AI tools enthusiast. I test, benchmark, and review AI models and developer tools so you don't have to.
187
Articles
103,517
Total Views
328K
Words Written
All Articles (187 total)
Install Ollama in 10 Minutes: Run Any LLM Locally Free
A no-fluff walkthrough for installing Ollama on Mac, Windows, or Linux and running open models like Llama 3.3 and DeepSeek V3 completely offline.
Gemini 3.1 Pro vs 2.5 Pro: 6 Real Upgrades
A data-driven comparison of Gemini 3.1 Pro vs Gemini 2.5 Pro: benchmarks, pricing, agentic tools, and whether the upgrade is actually worth it in late 2026.
Local LLMs vs Cloud APIs: Cost, Privacy, and Performance
A practical comparison of local LLMs versus cloud APIs from OpenAI, Anthropic, and Google—covering real-world costs, privacy trade-offs, inference performance, and when each approach makes sense.
Pocket LLM Benchmarks: What Actually Runs on Your Phone
The Artificial Analysis mobile inference data shows a widening gap between what phones can theoretically run and what they can sustain. Small models won, decoding is slow, and thermals still rule everything.
Llama API Tutorial: Build Your First App in 30 Minutes
A no-fluff Llama API tutorial that ships a working streaming chatbot in under 30 minutes. Real code, real keys, real cost numbers, no GPU required.
Run Mistral Large 3 Locally: GPU Setup & Real Benchmarks
A practical guide to running Mistral Large 3 on your own server hardware. VRAM math for the 675B MoE, vLLM setup, and what to expect from FP8 and NVFP4 deployments.
DeepSeek V4-Pro Review: The Open-Source Reasoning King?
A candid look at DeepSeek V4-Pro for reasoning workloads, covering benchmarks, pricing, real workflow trade-offs, and whether it can dethrone Claude and GPT-5 in 2026.
LoRA Speedrun: Wall-Clock Leaderboard Shaking Up Fine-Tuning
A new public leaderboard called LoRA Speedrun ranks fine-tuning techniques by raw wall-clock time. The results expose which tricks actually matter, and which are just marketing.
Choosing an AI Model: 11 LLMs, One Prompt, Wildly Different
One prompt, 11 top AI models, wildly different outputs. Here's an opinionated breakdown of which LLM to pick for coding, reasoning, writing, and long-context work in 2026.
Grok 4.5 vs Grok 4: 7 Real Differences That Matter
A no-fluff breakdown of Grok 4.5 vs Grok 4, from reasoning gains and coding speed to context, pricing, and whether the upgrade is actually worth it.
Run Qwen 3.6-27B Locally: 5-Step GPU Setup Guide
A practical tutorial for running Qwen 3.6-27B on your own GPU. Includes AWQ setup, vLLM serving, real tokens-per-second benchmarks, and the pitfalls that eat hours.
Gemini 3.5 Pro Review: 7 Reasoning Tests That Matter
An honest Gemini 3.5 Pro review focused on reasoning: benchmarks, pricing, real-world tradeoffs, and whether Google's mid-2026 flagship is worth switching to.