
Google has officially unveiled Gemini 4 Argon, its latest frontier AI model, promising major improvements in coding, long-running tasks, professional knowledge work and cybersecurity. But unlike a typical Gemini or AI model launch, you probably won’t be able to try this one immediately.
Announced by Google earlier today, Gemini 4 Argon is initially being made available to a selected group of trusted cybersecurity defenders through Google’s Fairwind Program. Wider availability for developers, businesses and consumers will come later, starting with paid API customers and Google AI Ultra subscribers. Google hasn’t given us an exact date yet.
And I can understand why the company is taking a more cautious approach this time.
Gemini 4 Argon can keep working for much longer
One of Argon’s most eye-catching improvements is its enormous 1 million-token output limit, up from 64,000 tokens previously. Google says the extra headroom allows Argon to reason through extremely long and complicated tasks without running out of room halfway through them. Its input context window is also one million tokens, according to early technical analysis.
This feels particularly relevant as AI increasingly moves from answering individual prompts towards performing longer jobs autonomously. Google says thousands of its own employees are already using Argon internally. Some of the examples are considerably more ambitious than asking Gemini to rewrite an email.
Argon agents have been involved in migrating C and C++ code to Rust, including work involving more than 800,000 lines of code in the Fuchsia Zircon kernel. Another group of agents analysed Google’s data-centre telemetry and identified memory optimisations that reportedly freed more than 300 TiB of memory after deployment. Google also claims Argon helped improve a quantum computing algorithm beyond a previously published baseline.
Obviously, these are Google’s own examples, so I’d still like to see how Argon behaves when developers outside Google start throwing messy real-world projects at it.
Google has specifically trained Gemini 4 Argon for defensive cybersecurity work. The company says the model can autonomously find, validate and patch critical software vulnerabilities.
On CWE-bench v1, Argon scored 68%, tying for the highest score according to Google’s results. Security company Wiz has also been testing Argon through its Scan for Good initiative, with Google saying the model discovered a critical vulnerability affecting healthcare software that previous frontier models had missed.
But these abilities are also part of the reason Google isn’t simply switching Argon on inside Gemini for everyone today. Google says it is strengthening safeguards against malicious cyber use, CBRN-related misuse, indirect prompt injection attacks and models taking actions beyond a user’s intentions. It’s also testing systems that monitor Argon’s reasoning and actions and can stop execution when necessary.
We’re reaching a point where the headline feature of a new AI model isn’t simply that it can write or reason better, but that companies are increasingly thinking about what happens when these models can independently perform complicated technical tasks for long periods of time.
Google’s benchmarks naturally paint a very impressive picture. Argon scores 77.9% on DeepSWE v1.1 for long-horizon software engineering, tops the Vals Index covering professional work such as finance, coding, legal and tax tasks, and scores 91.7% on LVBench for long-video understanding.
Independent testing gives us a little more perspective. Artificial Analysis reportedly scores Gemini 4 Argon at 53 on its Intelligence Index, putting it roughly alongside other current frontier models rather than clearly ahead of everything else. Interestingly, Argon appears to use substantially more output tokens to complete some tasks, so that giant output allowance may be doing quite a bit of work.
Google will initially price Gemini 4 Argon at $2 per million input tokens and $10 per million output tokens, before eventually increasing those rates to $4 and $20 respectively.
Gemini 4 Argon looks like an important step towards AI models that don’t just answer harder questions, but can stay focused on complicated jobs for much longer. The bigger question is when Google will feel comfortable enough to let the rest of us use it.






