Skip to content

Google's Gemini 4 Argon Goes To Cyber Defenders First

Google's new frontier model claims the lead on coding, legal and finance benchmarks, but security teams get it before developers, a sign of how cautiously the top AI labs now release models.

Google's Gemini 4 Argon Goes To Cyber Defenders First
Image courtesy: Unsplash

Google has unveiled Gemini 4 Argon, its first new flagship model since the Gemini 3 series last November, and it is not letting most people use it yet. The company said on 30 September that Argon is rolling out first to cyber defenders in its Fairwind Program, with paying API customers and Google AI Ultra subscribers next in line, although it has not given a date.

Koray Kavukcuoglu, who leads Google DeepMind and serves as Google's chief AI architect, said releasing a model with Argon's capabilities "requires a phased approach." Google has also given the model to the US government under its voluntary pre-release access process, and thousands of Google employees already use it for coding, research and writing.

Google aims Argon at long, multi-step work in three areas: software engineering, professional knowledge work in fields such as law and finance, and cybersecurity. The model can now produce up to 1 million tokens of output in one go, up from 64,000 in earlier Gemini models, which lets it write or rewrite very large bodies of code or documents in a single task.

Where Argon Claims The Lead

On Google's own published results, Argon leads or ties on 13 of 18 benchmarks against OpenAI's GPT-6 Astra and Anthropic's Claude Opus 5.5, according to VentureBeat's count. The largest gaps come in business tasks rather than pure reasoning.

On Harvey's legal agent benchmark, which tests legal research and drafting, Argon scored 19.6% against 5.4% for GPT-6 Astra and 3.8% for Claude Opus 5.5. On AutomationBench, which measures how well a model completes business processes from start to finish, it scored 51.3% against 42.5% for Claude Opus 5.5 and 41.4% for GPT-6 Astra, and on the Vals Finance Agent test for multi-step financial research it reached 65.4%, about seven points clear of the next model.

In software engineering, Argon scored 77.9% on DeepSWE v1.1, a few points ahead of both rivals. It also topped LVBench, a test of understanding long videos, with 91.7%.

Where It Still Trails

Argon does not win everywhere. GPT-6 Astra beat it by more than ten points on FrontierSWE v2 and on Terminal-Bench Science, and Claude Opus 5.5 led on Terminal-bench 4.0, which tests how well a model works through tasks in a command-line environment.

All of these figures come from Google, and benchmark wins rarely translate directly into performance on a company's own work. Independent evaluations, and the experience of early enterprise users, will give a clearer picture once the model reaches paying customers.

Why Security Teams Get It First

The cyber-first release says as much about the state of frontier AI as the benchmark tables do. Models that can find and fix software flaws on their own can also find flaws for attackers to exploit, and the major labs have moved over the past six months to put their most capable cyber tools in defenders' hands before anyone else's.

Anthropic set the pattern in April with Project Glasswing, which gave a restricted model to a group of large technology and security companies, and OpenAI followed in May with Daybreak, a tiered programme that gives verified defenders access to more capable cyber tools. Google launched its own Fairwind Program on 2 September, initially built around a smaller model, Gemini 3.8 Flash Cyber, and a tool called CodeMender that finds, verifies and fixes vulnerabilities.

More than 650 organisations have joined Fairwind, including CrowdStrike, Palo Alto Networks, Snowflake and Wiz, alongside government agencies and operators of critical infrastructure in healthcare, telecoms, energy and finance. Participants must limit access to their internal security, incident response and penetration testing teams. Google says members can generate "verified, deployment-ready patches in minutes" rather than spending weeks on manual fixes.

The Washington Factor

The release also follows a new US framework. A June executive order set up a voluntary process in which developers of the most capable models can give the government access for up to 30 days before a wider launch, followed by release to trusted partners chosen together with officials. The process places no licensing requirement on the labs, but it has quickly become part of how frontier models reach the market.

What Argon Has Already Done Inside Google

Google backed its claims with examples from its own systems. It said Argon-powered agents freed more than 300 tebibytes of memory across its data-centre fleet through software optimisation, and migrated large C and C++ projects to Rust, a language designed to avoid the memory errors behind many security flaws, including parts of the kernel of its Fuchsia operating system.

In one case, agents rewrote libgav1, a video decoder, and replaced 32,000 lines of low-level code with a Rust version that ran 2.7 times faster than an earlier Rust port. Outside Google, the cloud security company Wiz used Argon in its Scan for Good initiative and found a critical vulnerability exposing personal data in healthcare software used by hospitals worldwide, a flaw Google says earlier frontier models had missed.

Why This Matters Beyond The Data Centre

The C and C++ to Rust work matters well beyond Google's servers. Much of the software inside connected devices, from routers and industrial controllers to medical equipment, still runs on C and C++ firmware that is costly to audit and rewrite. If models like Argon can carry out those migrations reliably and at scale, device makers could cut a whole class of memory-safety flaws from their products, which grows more important as AI moves from screens into machines, a shift we traced in how physical AI turns IoT sensing into action.

Pricing Aimed At Enterprise Adoption

Google has priced Argon to win business customers when it opens up. During an introductory period, the API will cost $2 per million input tokens and $10 per million output tokens, rising later to $4 and $20, with cached input tokens discounted by 95%. At the standard rate, that matches Claude Opus 5.5 and sits well below GPT-6 Astra, which VentureBeat put at $10 and $50.

The introductory pricing will make Argon one of the cheaper frontier options for companies testing long-running agents, where output volumes and cached context can drive costs up quickly. The 1 million token output limit adds to the appeal for tasks such as large code migrations or long document drafting, although companies will need to watch how quickly long outputs add to their bills.

Safety Built Into The Launch

Google has paired the release with safety measures across four areas: defences against misuse for cyber or chemical, biological, radiological and nuclear attacks; resistance to prompt injection, where hidden instructions in documents or web pages hijack an AI agent; monitoring of the model's reasoning for signs of misaligned behaviour; and hardened, sandboxed environments for high-risk testing.

Argon recorded a 0.7% attack success rate on the Gray Swan indirect prompt injection test, against 1.0% for Claude Opus 5.5 and 8.5% for GPT-6 Astra. "We strongly encourage the rest of the industry to preserve reasoning transparency in these pivotal moments," Google said, arguing that the ability to read a model's reasoning is a key safeguard as agents gain autonomy.

What Business Buyers Should Weigh

For companies choosing an AI platform, Argon adds a strong option for agentic work in software, legal and finance, but it is not yet one they can use. Most buyers will have to wait for the API release, and the gap between Google's benchmarks and real workloads remains to be tested.

Security leaders in critical sectors have the clearest opportunity now, since Fairwind membership gives early access to tools that can find and patch flaws faster than manual teams. For everyone else, the more lasting change may be in how frontier models arrive: through government review, then defenders, then paying customers, rather than in a single public launch.

Frontier AI Now Arrives In Stages

Gemini 4 Argon puts Google back near the front of the frontier model race after ten months without a new flagship, with clear leads on the business tasks that enterprises pay for and gaps that rivals can still point to. The way Google released it may outlast the benchmark scores, because staged access, starting with cyber defenders and government reviewers, has become the norm for the most capable models across all three leading US labs.

The real test comes when paying customers can put Argon to work on their own code and documents. If its results hold up outside Google's tables, and its pricing stays as low as the introductory rate suggests, the frontier race will turn less on who leads a benchmark and more on who can deliver capable agents safely and at a cost businesses can sustain.

Add Morning Tick on Google