Skip to content

Gemini

(30 articles)

Vals AI: The a16z-Backed Push for Honest Benchmarks

Andreessen Horowitz is betting on Vals AI to fix a broken benchmarking scene. Here's what the platform does differently, and why neutral testing matters more...

September 19, 20268 min

Gemini 3.1 Pro Review: 7 Reasoning Tests, One Verdict

An honest Gemini 3.1 Pro review focused on reasoning. Benchmark scores, real-world use cases, pricing, and whether Google finally caught Claude and GPT.

September 17, 20268 min

Gemini 3.5 Flash-Lite vs 2.5 Flash-Lite: 7 Real Changes

What actually moved between Gemini 2.5 Flash-Lite and 3.5 Flash-Lite: pricing, latency, benchmarks, and whether the upgrade is worth it (spoiler: 3.5...

September 16, 202610 min

Claude Opus 4.6 vs Gemini: The Honest 2026 Verdict

Gemini just landed on Windows and Claude Opus 4.6 is still charging premium rates. Which one actually deserves your subscription? A data-driven breakdown of...

September 12, 202610 min

Gemini Omni 1.1 Flash vs 1.0: 7 Real Changes

A no-fluff breakdown of what Gemini Omni 1.1 Flash actually improves over Gemini Omni 1.0, from latency to pricing to multimodal quality. Verdict included.

September 3, 20269 min

Gemini 3.1 Pro vs 2.5 Pro: 6 Real Upgrades

A data-driven comparison of Gemini 3.1 Pro vs Gemini 2.5 Pro: benchmarks, pricing, agentic tools, and whether the upgrade is actually worth it in late 2026.

August 31, 20269 min

Choosing an AI Model: 11 LLMs, One Prompt, Wildly Different

One prompt, 11 top AI models, wildly different outputs. Here's an opinionated breakdown of which LLM to pick for coding, reasoning, writing, and long-context...

August 31, 202611 min

Gemini 3.5 Pro Review: 7 Reasoning Tests That Matter

An honest Gemini 3.5 Pro review focused on reasoning: benchmarks, pricing, real-world tradeoffs, and whether Google's mid-2026 flagship is worth switching to.

August 25, 20269 min

Gemini 3.5 Pro vs Claude Fable 5: Best Agent Model in 2026

Two frontier models, one agentic crown. We break down Gemini 3.5 Pro vs Claude Fable 5 on tool use, SWE-bench, pricing, and long-horizon reliability.

August 18, 20269 min

LLM Agents Flop at Coordination: Inside the ALEM Benchmark

A new open-ended coordination benchmark tests 13 LLMs across communication, trading, crafting, and combat. Most agents average just 6% normalised return.

July 19, 20268 min

Stop Google Training AI on You: 7 Settings to Fix Now

Google quietly expanded which of your data can train Gemini. Walk through the 7 exact toggles that pull your account back out of the training pool.

July 12, 20267 min

5 Google Search Hacks That Crush Thrift & Vintage Hunting

Google quietly rolled out AI features that turn random thrift hauls into curated vintage scores. Five ways to use Search, Lens, and Shopping to find the good...

June 5, 20268 min
Page 1 of 3Next