Gemini 4 Pro Leaked: What the Chart Says, and One Prompt to Test It Against Claude on Your Own Work
A leaked chart puts Google's next model ahead of Claude Opus 5.5 and GPT-6 Astra on a top coding test, at about half Claude's price. What it shows, what it doesn't, and the one prompt to paste into all three the day it ships.
A chart posted on X on September 28 shows a model called Gemini 4 Pro, Google's next AI, beating Claude Opus 5.5 and OpenAI's GPT-6 Astra on DeepSWE, a hard coding test: 88.7 against 74.2 and 74.1. The same chart prices it at about half of what Claude costs and gives it a 2 million token context window, twice Claude's.
Google hasn't announced Gemini 4 Pro, and nobody at Google has confirmed the chart. So treat every number below as a leak until the model ships. The post is here: https://x.com/rynorhn/status/2104497579463786839
What the leaked chart says
| Benchmark | Gemini 4 Pro (leak) | Claude Opus 5.5 | GPT-6 Astra |
|---|---|---|---|
| DeepSWE v1.1 (coding) | 88.7 | 74.2 | 74.1 |
| Terminal-Bench 2.1 | 95.3 | about 88 | about 88 |
| OSWorld 2.0 (using a computer) | 86.8 | 81.8 (partial) | 72.6 (offline, partial) |
| GDPval-AA (work tasks) | 2,064 (v2) | 1,846 (v2.1) | 1,542 (v2.1) |
| Context window | 2M tokens | 1M tokens | 1.1M tokens |
| Price per 1M tokens (in / out) | $2.25 / $11.25 | $4 / $20 | $10 / $50 |
Read it the way the chart itself asks you to. On DeepSWE the gap is big: about 15 points. Claude and Astra are 0.1 apart there, which is a tie. The price is 56% of Claude's for both input and output, so about half. Some cells come with notes: OSWorld is marked partial, and the GDPval-AA scores come from different versions of the test (v2 against v2.1), so those two rows don't compare like for like.
Keep reading
Free, for an email.
Unlocks 6 more sections, 2 copy-paste prompts, the document version, and every other guide on the site, for free. Enter your email once to keep reading.
19,000+ follow where these guides come from.No spam. Unsubscribe in one click.