Back to list
Oct 01, 2026
174
0
0
GeminiNEW

Gemini 4 Argon: Frontier Model Opens to Cyber Defenders

Google announced Gemini 4 Argon on Sept 30, giving trusted cyber defenders first access. Price is $2/$10 per 1M tokens; broad release has no date.

#Gemini 4 Argon#Google DeepMind#Gemini#Fairwind Program#Cybersecurity
Gemini 4 Argon: Frontier Model Opens to Cyber Defenders
AI Summary

Google announced Gemini 4 Argon on Sept 30, giving trusted cyber defenders first access. Price is $2/$10 per 1M tokens; broad release has no date.

Introduction

Google announced Gemini 4 Argon on September 30, 2026, describing it in its official blog post as "our new frontier model." The announcement carries a catch that the headlines treated differently. TechCrunch titled its story "Google releases Gemini 4 Argon," while Ars Technica titled its own "you can't use it yet." Both are consistent with Google's text. Argon is "rolling out to a set of trusted cyber defenders through our Fairwind Program," and Google says it will make the model available to developers, enterprises, and consumers "as soon as possible," with no date attached. The post positions Argon for real-world software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense.

Feature Overview

Who can use it today. Per Google's post, access is limited to a set of trusted cyber defenders in the Fairwind Program, plus Google's own internal teams. Google also writes that it is "actively engaged in the U.S. government's voluntary process for pre-release model access while we gradually expand access." When broader availability arrives, the post says it will start with paid API customers and Google AI Ultra subscribers. The Verge reported that Google is limiting access at first so it can make sure the model is not misaligned, and quoted Google DeepMind SVP Koray Kavukcuoglu on the phased approach.

Pricing and output limit. Google lists an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off the input price. A footnote in the post says that after the introductory period expires, $4 per million input tokens and $20 per million output tokens will apply; Google does not say how long the introductory period lasts. The output token limit rises to 1 million tokens, up from the previous 64K, which Google calls "industry-leading."

Reported benchmarks (all figures from Google's post).

BenchmarkWhat it measuresArgon result
DeepSWE v1.1Real-world long-horizon software engineering77.9%, described as a new state of the art
CWE-bench v1Remediating security vulnerabilities68%, tied for first place
AutomationBench (Zapier)End-to-end execution across business functions51.3%, ranked #1
LVBenchLong video understanding91.7%, described as state of the art
Vals IndexEconomic impact across finance, coding, legal, and tax workDescribed as the leading model, no score given in the text

Google's blog also presents a comparison chart against competing models. Ars Technica read that chart as showing Argon's DeepSWE v1.1 score ahead of GPT-6 Astra, Fable 5.1, and Opus 5.5, and TechCrunch relayed Google's claim of higher scores than OpenAI's and Anthropic's models. These are vendor-reported figures, and the text of the post does not describe independent replication.

Cybersecurity capability. Google says it trained Argon to "autonomously find, validate, and patch critical software vulnerabilities," and that trusted defenders and Google's internal teams will receive it "without cyber guardrails." Security firm Wiz is already using it through its Scan for Good initiative, and Google says the model found a critical vulnerability exposing personal information across healthcare software used by hospitals worldwide that "previous frontier models had missed." Against 3.8 Flash Cyber, Google says Argon outperforms it on Wiz's internal black-box penetration testing benchmark, though the post gives no numbers for that test. Ars Technica noted that Google did not provide specifics on the healthcare flaw.

Internal use at Google. The post lists several examples: a quantum-computing subroutine optimization that beat a published baseline by 40%; agents that identified memory optimizations freeing over 300 TiB across data centers once rolled out, with an estimated total of 500 TiB to 1 PiB; and C/C++ to Rust migrations that scale up to 800K+ lines for the Fuchsia Zircon kernel. For the libgav1 video decoder, Google says agents replaced 32K lines of SIMD code and produced a decoder that runs 2.7x faster than the existing Rust port with identical video output.

Usability Analysis

Almost nobody reading this can run Argon today, so usability is a question of access policy rather than product experience. The audience that can use it is vetted security teams, who get a variant without cyber guardrails. Everyone else is waiting on a rollout that Google has sequenced but not dated.

Google describes four safeguard areas it is strengthening before broad release: defending against misuse (including cyber and CBRN attacks, with monitoring of the model's internal activations), prompt injection resistance, monitoring for misalignment by watching the chain-of-thought and actions, and hardening its sandboxed environments. It calls Argon its "most resilient model yet" against indirect prompt injection and says it leads on Gray Swan's Indirect Prompt Injection benchmark, without publishing a score in the text.

For developers planning ahead, the published prices and the 1M output limit are the concrete planning inputs. Because the introductory period has no stated end date, budgets are safer modeled at the $4/$20 post-introductory rate.

Pros and Cons

The strengths are a large output window for long-running tasks, a published price, a broad set of reported benchmark results, and a safeguards-first release sequence. The limits are the lack of general availability or a timeline, benchmark claims that so far come from Google itself, and the dual-use nature of a guardrail-free security model, even in vetted hands.

Outlook

Google's own sequencing suggests the next milestone is expanded access as feedback from the initial cohort shapes the guardrails, followed by paid API customers and Google AI Ultra subscribers. The post does not commit to dates, so any schedule would be speculation. Two things to watch are independent evaluations once outside researchers get access, and how the U.S. government's voluntary pre-release process shapes the timing Google describes.

Conclusion

Gemini 4 Argon is best read as a staged launch: a capability announcement and price now, restricted access for security defenders, and broad availability later. Its reported results in software engineering and vulnerability remediation are strong on Google's own telling, but they remain vendor-reported until outside testers confirm them. Security teams and enterprise planners should track the Fairwind Program and Google's availability updates. Rating: 4 out of 5, with the caveat that most users cannot yet try it.

Editor's Verdict

Gemini 4 Argon: Frontier Model Opens to Cyber Defenders earns a solid recommendation within the Gemini space.

The strongest case for paying attention: output limit of 1M tokens supports long, multi-step agentic tasks in a single trajectory. That alone raises the bar for what readers should expect in this space. Reinforcing that, pricing is published up front at $2 input and $10 output per million tokens, with a 95% discount on cached input, rising to $4 and $20 after the introductory period — practical value rather than just headline appeal. The broader signal worth registering is straightforward: staged access separates the capability announcement from availability, so the headline 'release' and the 'you can't use it yet' framing are both accurate. On the other side of the ledger, one constraint is real rather than a marketing footnote: general availability has no date, and only trusted cyber defenders and Google's internal teams can use the model today. It should factor into any serious decision. Layered on top of that, most benchmark results come from Google's own post, and the Vals Index and prompt injection claims come without scores in the text — which narrows the set of teams for whom this is an obvious yes.

For Google Cloud and Workspace integrators, multimodal-first teams, and Gemini API adopters, this is a model to plan around now and evaluate as soon as access opens beyond the Fairwind Program. For everyone else, the safer posture is to monitor coverage and revisit once the use cases that matter to your team are demonstrated in the wild.

Advertisement

Pros

  • Output limit of 1M tokens supports long, multi-step agentic tasks in a single trajectory.
  • Pricing is published up front at $2 input and $10 output per million tokens, with a 95% discount on cached input.
  • Reported results span software engineering, vulnerability remediation, business automation, and long video understanding.
  • Release plan puts safeguards and vetted users first before wider availability.

Cons

  • General availability has no date, and only trusted cyber defenders and Google's internal teams can use the model today.
  • Most benchmark results come from Google's own post, and the Vals Index and prompt injection claims come without scores in the text.
  • A variant without cyber guardrails carries dual-use risk, even though it is restricted to vetted defenders.
  • Post-introductory pricing doubles to $4 input and $20 output per million tokens, and Google does not say when the introductory period ends.
Advertisement

Comments0

Key Features

1. Access: rolling out to trusted cyber defenders through the Fairwind Program and Google's internal teams; broader release to developers, enterprises, and consumers is planned without a stated date. 2. Pricing: introductory $2 per million input tokens and $10 per million output tokens, with cached input tokens at 95% off; $4 / $20 applies after the introductory period. 3. Output limit: raised to 1M tokens from 64K. 4. Reported results: DeepSWE v1.1 77.9%, CWE-bench v1 68% (tied for first), AutomationBench 51.3% (#1), LVBench 91.7%. 5. Cyber focus: trained to find, validate, and patch vulnerabilities; defenders receive a variant without cyber guardrails.

Key Insights

  • Staged access separates the capability announcement from availability, so the headline 'release' and the 'you can't use it yet' framing are both accurate.
  • Sequencing starts with vetted security defenders, which signals that Google treats offensive-capable cyber skill as the gating risk for this model.
  • Output length is a quiet headline change, since a 1M-token output limit up from 64K lets one trajectory run far longer before stopping.
  • Benchmark claims are vendor-reported, so the numbers are best treated as Google's own evidence until outside evaluators run the model.
  • Pricing is published in two stages, $2/$10 at launch and $4/$20 afterwards, and with no end date for the introductory period, production budgets are safer modeled at the higher rate.
  • Safeguards are described in four areas, including monitoring of internal activations and chain-of-thought, which shows how access policy and model monitoring are being bundled together.
  • Internal examples such as the libgav1 rewrite and the data-center memory savings show the use cases Google chose to demonstrate, not a guarantee of results for outside users.

Was this review helpful?

Share

Twitter/X
Advertisement