
Gemini 4 Argon arrives with top benchmark scores and limited access
Google has unveiled Gemini 4 Argon, its new flagship model, claiming top results against OpenAI and Anthropic on most benchmarks while access stays limited to a small group of cybersecurity partners.
Google announced Gemini 4 Argon on Wednesday, presenting the model as its new flagship and keeping it limited for now to a small group of cybersecurity partners. The company says it is taking a phased approach and has engaged in the U.S. government's voluntary pre-release access process while it gradually expands availability.
Benchmarks favour knowledge work
Across Google's own test suite, Argon takes first place in 13 of 18 benchmarks against Anthropic's Claude Opus 5.5 and Fable 5.1 and OpenAI's GPT-6 Astra. The gap is widest on tasks the company groups as knowledge work: Argon scores 68.9% on the Vals Index, against 63.1% for GPT-6 Astra, and 65.4% on the Vals Finance Agent v2, where the nearest rival reaches 58.9%. On Zapier's AutomationBench it records 51.3%, close to nine points ahead of Opus 5.5.
Coding results are mixed
The coding picture is less clear. Google highlights a new state of the art of 77.9% on DeepSWE v1.1, yet Argon finishes last on both FrontierSWE v2 and Terminal-Bench 4.0, where GPT-6 Astra and Opus 5.5 lead it by 10.5 and nine points. It scores 91.9% on Vibe Code Bench, a result the company notes sits alongside three rivals that all clear 89%. Google also reports 84.2% on the GraphWalks test for inputs between 256K and 1M tokens, more than 12 points ahead of GPT-6 Astra, and 19.6% on Harvey's Legal Agent Benchmark, nearly triple Fable 5.1 but still only about one fully completed task in five.
A million output tokens and fewer guardrails
One feature stands out where rivals are not pushing as hard: Argon can generate up to one million output tokens, up from 64,000 in earlier Gemini models. Google argues the added headroom lets the model reason through long problems in a single run. In cybersecurity, the company says it trained Argon to find, validate and patch vulnerabilities autonomously, and for Fairwind participants and its own teams it ships the model without cyber guardrails. Wiz, acquired by Google for $32 billion in March, is already using it in the Scan for Good initiative, where the company says the model found a critical flaw in healthcare software used by hospitals worldwide.
Access and pricing
Argon is effectively the new Pro model Google first outlined at its I/O conference in May. The original June launch slipped and produced a run of Flash models instead. Google says it will collect feedback from early testers and adjust guardrails before opening access beyond the Fairwind Program, after which API customers and AI Ultra subscribers go first. Pricing has not been announced. The announcement came a day after CEO Sundar Pichai co-signed a commitment to self-police AI development following a meeting with President Donald Trump, alongside Anthropic, Meta, Nvidia, OpenAI and SpaceX.
Sources: The New Stack · Techmeme
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.