Skip to content
S

Shadman Ahmed

Software Architect

Software architect and AI tools enthusiast. I test, benchmark, and review AI models and developer tools so you don't have to.

187

Articles

103,517

Total Views

328K

Words Written

All Articles (187 total)

Install Ollama in 10 Minutes: Run Any LLM Locally Free

A no-fluff walkthrough for installing Ollama on Mac, Windows, or Linux and running open models like Llama 3.3 and DeepSeek V3 completely offline.

September 2, 2026 16 min 113tutorials

Gemini 3.1 Pro vs 2.5 Pro: 6 Real Upgrades

A data-driven comparison of Gemini 3.1 Pro vs Gemini 2.5 Pro: benchmarks, pricing, agentic tools, and whether the upgrade is actually worth it in late 2026.

August 31, 2026 9 min 72comparisons

Local LLMs vs Cloud APIs: Cost, Privacy, and Performance

A practical comparison of local LLMs versus cloud APIs from OpenAI, Anthropic, and Google—covering real-world costs, privacy trade-offs, inference performance, and when each approach makes sense.

August 31, 2026 6 min 67comparisons

Pocket LLM Benchmarks: What Actually Runs on Your Phone

The Artificial Analysis mobile inference data shows a widening gap between what phones can theoretically run and what they can sustain. Small models won, decoding is slow, and thermals still rule everything.

August 31, 2026 7 min 103benchmarks

Llama API Tutorial: Build Your First App in 30 Minutes

A no-fluff Llama API tutorial that ships a working streaming chatbot in under 30 minutes. Real code, real keys, real cost numbers, no GPU required.

August 31, 2026 13 min 56tutorials

Run Mistral Large 3 Locally: GPU Setup & Real Benchmarks

A practical guide to running Mistral Large 3 on your own server hardware. VRAM math for the 675B MoE, vLLM setup, and what to expect from FP8 and NVFP4 deployments.

August 31, 2026 13 min 53tutorials

DeepSeek V4-Pro Review: The Open-Source Reasoning King?

A candid look at DeepSeek V4-Pro for reasoning workloads, covering benchmarks, pricing, real workflow trade-offs, and whether it can dethrone Claude and GPT-5 in 2026.

August 31, 2026 9 min 71reviews

LoRA Speedrun: Wall-Clock Leaderboard Shaking Up Fine-Tuning

A new public leaderboard called LoRA Speedrun ranks fine-tuning techniques by raw wall-clock time. The results expose which tricks actually matter, and which are just marketing.

August 31, 2026 8 min 44benchmarks

Choosing an AI Model: 11 LLMs, One Prompt, Wildly Different

One prompt, 11 top AI models, wildly different outputs. Here's an opinionated breakdown of which LLM to pick for coding, reasoning, writing, and long-context work in 2026.

August 31, 2026 11 min 53listicles

Grok 4.5 vs Grok 4: 7 Real Differences That Matter

A no-fluff breakdown of Grok 4.5 vs Grok 4, from reasoning gains and coding speed to context, pricing, and whether the upgrade is actually worth it.

August 27, 2026 9 min 120comparisons

Run Qwen 3.6-27B Locally: 5-Step GPU Setup Guide

A practical tutorial for running Qwen 3.6-27B on your own GPU. Includes AWQ setup, vLLM serving, real tokens-per-second benchmarks, and the pitfalls that eat hours.

August 26, 2026 13 min 268tutorials

Gemini 3.5 Pro Review: 7 Reasoning Tests That Matter

An honest Gemini 3.5 Pro review focused on reasoning: benchmarks, pricing, real-world tradeoffs, and whether Google's mid-2026 flagship is worth switching to.

August 25, 2026 9 min 97reviews
PreviousPage 2 of 16Next