Skip to content

Google Ships Gemini 4 Argon, Gates Autonomous Vulnerability Patching to a Pilot Program

· by Pondero Newsdesk

The short version

Google's new frontier model can rewrite C and C++ codebases into Rust at scale and autonomously patch critical vulnerabilities, but it ships only to vetted Fairwind Program partners, with no public pricing tier live yet.

Google Ships Gemini 4 Argon, Gates Autonomous Vulnerability Patching to a Pilot Program

Gemini 4 Argon can rewrite memory-unsafe C and C++ into Rust across million-line codebases, and a security partner says it already caught a critical hospital-software vulnerability that earlier frontier models missed. Google released Argon on September 30 as a frontier model built around autonomous vulnerability discovery and patching, but it is not available to the public: access runs through a small cohort of vetted cyber defenders in Google's Fairwind Program, per Google's announcement.

What Argon actually does

Google trained Argon to find, validate, and patch critical software vulnerabilities on its own, a capability the company is releasing to trusted defenders without the cyber guardrails applied to its general release, according to the announcement. The model's maximum output jumped to 1 million tokens, up from 64,000 tokens on the prior Gemini generation, a 16x increase Google says lets Argon reason through a single long, multi-step problem without breaking it into smaller calls. On Google's internal codebase-migration work, Argon agents are porting C and C++ libraries to Rust at a scale that includes the Fuchsia Zircon kernel, over 800,000 lines of code; for the video codec library libgav1, Argon replaced 32,000 lines of hand-tuned SIMD code and produced a Rust decoder that runs 2.7 times faster than the prior Rust port while matching the original video output exactly.

On benchmarks, Google says Argon sets a new state of the art on DeepSWE v1.1, a long-horizon software engineering test, with a score of 77.9%, and ranks first on Zapier's AutomationBench at 51.3%, a test of end-to-end business-task execution. It also leads the Vals Index, a cross-domain benchmark from AI evaluation startup Vals that weights finance, coding, legal, and tax performance by each sector's share of US GDP, and posts a state-of-the-art 91.7% on LVBench, a long-video-understanding test. On CWE-bench v1, a benchmark for patching known vulnerability classes, Argon ties for first place at 68%, building on the prior Gemini cybersecurity model's results on an earlier version of that benchmark. TechCrunch reported that Google's own framing goes further: the company says Argon scored significantly higher than OpenAI's GPT-6 Astra and Anthropic's Fable and Opus models across a range of benchmarks, a claim TechCrunch attributes to Google citing Vals's model index.

Pricing is set but not yet live for most users. Argon will launch at an introductory rate of $2 per million input tokens and $10 per million output tokens, with cached input priced 95% below standard input, before stepping up to $4 input and $20 output once the introductory window closes, per a footnote in Google's announcement.

Only vetted cyber partners can use Argon's vulnerability-hunting powers right now

Outside the Fairwind Program, Argon's headline capability is not something an AI-tool operator can test or buy today. Google says it is running Argon through the US government's voluntary pre-release access process and plans to open the model to API customers and Google AI Ultra subscribers only after that review and further guardrail testing conclude, with no date attached. Security vendor Wiz already has access through its Scan for Good initiative, a program that audits public infrastructure for free, and used Argon to uncover a vulnerability exposing sensitive patient data across hospital software used worldwide, a risk Google says earlier frontier models had missed. That single result is a concrete preview of what a wider release could mean for bug-bounty and penetration-testing workflows once it ships, but it is one disclosed case, not an independent audit. Teams planning around AI-assisted vulnerability management should treat Argon's benchmark wins as Google citing its own choice of third-party leaderboards, Vals and CWE-bench, rather than results an outside lab has reproduced, since nobody outside the pilot can run the model yet.

Context and reactions

Argon arrives into a frontier-model race that has moved fast this year. OpenAI released GPT-6 Astra on September 3 and called it its best model yet, and Anthropic shipped Claude Fable earlier this year with comparable claims, according to TechCrunch's reporting. Google frames Argon's benchmark lead against both of those models, though the comparison rests on Vals's scoring rather than a neutral third party running all three side by side. The release also lands as Google's broader AI push gains ground: the company said in August that its Gemini app had surpassed 1 billion monthly users, a milestone that puts it in the same range OpenAI reported for ChatGPT around the same time, per TechCrunch. Google was described as trailing in the AI race as recently as mid-2026; Argon's cybersecurity framing and internal productivity claims, including a reported 40% improvement over a published baseline in quantum-algorithm optimization and more than 300 TiB of memory freed across Google's data centers by Argon-run optimization agents, are the company's case that the gap has closed.

What to watch next

Watch whether Google widens the Fairwind Program or ships a general-availability API variant before the introductory pricing window ends, since that step change in price is the clearest signal of when broad access actually starts. Also watch for independent confirmation of Argon's cyber and coding benchmark claims from a source other than Vals or Google itself, given that CWE-bench and DeepSWE results currently come only from Google's own citation.

Sources