Prompting

Gemini 4 Pro Leaked: What the Chart Says, and One Prompt to Test It Against Claude on Your Own Work

A leaked chart puts Google's next model ahead of Claude Opus 5.5 and GPT-6 Astra on a top coding test, at about half Claude's price. What it shows, what it doesn't, and the one prompt to paste into all three the day it ships.

5min read
2prompts to copy
5numbered steps
7sections

A chart posted on X on September 28 shows a model called Gemini 4 Pro, Google's next AI, beating Claude Opus 5.5 and OpenAI's GPT-6 Astra on DeepSWE, a hard coding test: 88.7 against 74.2 and 74.1. The same chart prices it at about half of what Claude costs and gives it a 2 million token context window, twice Claude's.

Google hasn't announced Gemini 4 Pro, and nobody at Google has confirmed the chart. So treat every number below as a leak until the model ships. The post is here: https://x.com/rynorhn/status/2104497579463786839

What the leaked chart says

BenchmarkGemini 4 Pro (leak)Claude Opus 5.5GPT-6 Astra
DeepSWE v1.1 (coding)88.774.274.1
Terminal-Bench 2.195.3about 88about 88
OSWorld 2.0 (using a computer)86.881.8 (partial)72.6 (offline, partial)
GDPval-AA (work tasks)2,064 (v2)1,846 (v2.1)1,542 (v2.1)
Context window2M tokens1M tokens1.1M tokens
Price per 1M tokens (in / out)$2.25 / $11.25$4 / $20$10 / $50

Read it the way the chart itself asks you to. On DeepSWE the gap is big: about 15 points. Claude and Astra are 0.1 apart there, which is a tie. The price is 56% of Claude's for both input and output, so about half. Some cells come with notes: OSWorld is marked partial, and the GDPval-AA scores come from different versions of the test (v2 against v2.1), so those two rows don't compare like for like.

Keep reading

Free, for an email.

Unlocks 6 more sections, 2 copy-paste prompts, the document version, and every other guide on the site, for free. Enter your email once to keep reading.

19,000+ follow where these guides come from.No spam. Unsubscribe in one click.

Read next