The Story
Google has announced Gemini 4 Argon, its new frontier model, built for long-running workflows across software engineering, enterprise knowledge work and cybersecurity defence. It is rolling out first to a set of trusted cyber defenders through the company's Fairwind Program, with paid API customers and Google AI Ultra subscribers to follow. No date has been given for general availability.
Koray Kavukcuoglu, who runs Google DeepMind day to day, called it the company's next era of frontier intelligence. Sundar Pichai said on X that there had been a great deal of discussion about the model and he wanted to give an early look as soon as possible.
The output token limit rises to 1 million, from 64,000. Google's argument is that when a model has room to generate hundreds of thousands of tokens in a single run, it can reason more deeply and solve a hard problem in one attempt rather than in fragments.
On the benchmarks Google published, Argon scores 77.9 percent on DeepSWE v1.1, a test of long-horizon software engineering, against 74.2 percent for Claude Opus 5.5 and 74.1 percent for GPT-6 Astra. It ties for first on CWE-bench v1, which measures the ability to remediate security vulnerabilities, at 68 percent, and ranks first on Zapier's AutomationBench at 51.3 percent. Google also describes it as leading the Vals Index, which weights finance, coding, legal and tax work by economic contribution. The results are not a clean sweep, and on coding specifically it leads on two of four benchmarks.
Cybersecurity is the pillar Google has pushed hardest. It says Argon can autonomously find, validate and patch critical software vulnerabilities, and that it will be released without cyber guardrails to trusted defenders and internal Google teams so they can use its full defensive capability. Google says Wiz used Argon to find a critical flaw in healthcare software used by hospitals worldwide.
The company also points to internal use: helping its quantum computing researchers optimise subroutines, where it beat a published baseline by 40 percent in minutes, and producing a memory-safe video decoder running 2.7 times faster than the Rust port with identical output.
Bloomberg reported that while Gemini 4 has performed well on benchmarks, employees with access found it less impressive on real work and that it struggled with some coding tasks. Google said it would be inaccurate to say the model underperforms in areas such as coding, pointing to its own internal deployment.
Introductory pricing is $2 per million input tokens and $10 per million output tokens. Parameter count, architecture, training compute, latency and throughput have not been disclosed.
Pichai signed a voluntary AI safety accord with President Donald Trump and other major technology leaders this week.
Why It Matters
Frontier model launches have settled into a pattern, and this one breaks it in a specific way.
The usual sequence is a consumer release, a wave of screenshots, a benchmark table and an API a week later. Argon went to cybersecurity firms. Not as a marketing gesture, but as the first and for now only meaningful distribution, with the model's security restrictions deliberately removed for that audience. Consumers and most developers are waiting.
That ordering tells you what Google thinks it has built. This is not a model positioned on conversation or creativity; it is aimed at work that runs for hours, produces artefacts rather than replies, and is judged by whether the output compiles, holds up in court or closes a vulnerability. Software engineering, legal and financial analysis, and defensive security are all domains where the value of an AI system is measured against a professional's time rather than against another chatbot.
Cybersecurity is where that logic bites hardest, because the field has an asymmetry that favours attackers. A defender must find and fix every vulnerability in a codebase. An attacker needs one. Defenders are outnumbered, security engineers are scarce and expensive, and the volume of code the world runs on grows faster than anyone can audit it. A system that can read a large codebase, identify a flaw, verify it is real and write the patch attacks that asymmetry directly, which is why the million-token output limit and the security focus belong in the same announcement.
The same properties are what make it dangerous, and Google's handling reflects that. The skill of finding an exploitable bug does not distinguish between motives; only the user does. Giving it to defenders first, with the restrictions lifted only for them, is an attempt to get the capability into the hands of the people patching before it reaches those probing. Whether a release sequence can hold that line for long is the question the whole industry is now working through, and this is a fairly clear statement of one company's answer.
The Strategic Read
Two things in this launch deserve more attention than the benchmark table.
The first is the phrase "without cyber guardrails." Google is saying plainly that a version of this model exists which will help with security work that the restricted version refuses, and that access to it depends on being approved rather than on paying. That is a meaningful shift in how frontier models are distributed, and it follows from an uncomfortable symmetry: the capability to autonomously find, validate and patch a critical vulnerability is the same capability, pointed the other way, to autonomously find and exploit one. A model good enough to be genuinely useful to defenders is by construction good enough to be dangerous in other hands.
Releasing to defenders first is the reasonable response to that, and it also makes the distribution itself a safety mechanism rather than a commercial decision. The question it raises is who decides who is trusted. Fairwind is Google's programme, with Google's criteria, and a private company is now operating an access list for a capability with national-security implications. That is not a criticism of the choice so much as an observation that it is now normal, and that Pichai signing a voluntary safety accord with the US administration the same week is part of the same picture.
The second is the gap between the benchmarks and the people using it. Bloomberg's reporting that Google employees found the model less impressive on real work than on tests is the most important sentence in the coverage, and Google's denial does not settle it. Benchmarks are a proxy that models are increasingly optimised against, and the divergence between a leaderboard score and whether something is useful on a Tuesday afternoon has been widening across the industry. When the doubt comes from inside the company, from people with access, it is worth more than a score.
That is also why the limited release cuts both ways. Restricting access to trusted defenders is defensible on safety grounds, and it also means nobody outside Google can independently verify the claims for some time. The CWE-bench result, the Wiz vulnerability discovery, the quantum optimisation beating a baseline by 40 percent: all are vendor-reported, none reproducible from what has been published. The healthcare flaw in particular is an attributed example rather than a documented one.
The million-token output limit is the change most likely to matter commercially, and it is easy to underrate because it sounds like a specification bump. Sixty-four thousand tokens is roughly a long chapter. A million is a small codebase. The difference is not that the model writes more but that it can hold an entire piece of work in one continuous attempt rather than being chopped into segments that lose the thread between them. For the long-running professional workflows Google is targeting, that is closer to the actual constraint than raw intelligence has been for a while.
Pricing of $2 and $10 per million tokens is described as introductory, which is a word worth remembering when the comparison charts start circulating.
For daily, sharp analysis of the biggest moves in the Indian business and startup ecosystem, follow StartupFox.
