Gemini 4 Argon Just Beat Claude at Coding: The Cheat Sheet for Which AI to Use for What
Google's new model tops a real-world coding test the day it was announced, but you can't use it yet. What it scored, what it costs, when it opens up, and which AI to pick for each kind of job until then.
Google announced Gemini 4 Argon on September 30. On DeepSWE v1.1, a test built from real, long software engineering tasks, it scored 77.9%, ahead of Claude Opus 5.5 at 74.2% and OpenAI's GPT-6 Astra at 74.1%. Google calls that a new state of the art.
The catch: you can't use it yet. For now it goes to cyber defenders in Google's Fairwind Program and to the U.S. government's pre-release testing. Paid API customers and Google AI Ultra subscribers are next, with no date. So the useful question isn't whether to switch today. It's which AI to use for which job now, and where Argon fits when it opens.
What Google announced
| Gemini 4 Argon | Claude Opus 5.5 | GPT-6 Astra | |
|---|---|---|---|
| DeepSWE v1.1 (real coding tasks) | 77.9% | 74.2% | 74.1% |
| Price per 1M tokens (in / out) | $2 / $10 (intro), then $4 / $20 | $4 / $20 | $10 / $50 |
| Can you use it today? | No: Fairwind and U.S. gov first | Yes | Yes |
Google also reports Argon tied for first on CWE-bench v1 (68%), first on AutomationBench (51.3%) and 91.7% on LVBench, a test of understanding long videos. It can write up to 1 million tokens in one answer, up from 64K, and cached input is 95% off. Two things to read into the table: Claude and Astra are 0.1 apart on DeepSWE, which is a tie, and Argon's lead is 3.7 points on one test.
Keep reading
Free, for an email.
Unlocks 3 more sections, the document version, and every other guide on the site, for free. Enter your email once to keep reading.
19,000+ follow where these guides come from.No spam. Unsubscribe in one click.