hotAI

• 9 min read

Gemini 4 Argon’s safety case now hinges on 1 million tokens

Google’s Gemini 4 Argon pairs coding and cyber claims with a gated rollout, while its 1M-token output limit raises the burden on its safeguards.

Gemini 4 Argon’s safety case now hinges on 1 million tokens

Image: The Verge

Google has put Gemini 4 Argon behind a narrow access gate while claiming the model is its strongest system for software engineering, enterprise work, and defensive cybersecurity. Argon’s defining technical change is a 1 million-token output limit, up from 64,000 tokens in earlier Gemini models. Google says that capability requires stronger controls before developers or consumers can use it.

On September 30, 2026, Google began rolling out Argon through its Fairwind Program to a selected group of trusted cyber defenders. Broader access has no date. Google says it is participating in the U.S. government’s voluntary pre-release model-access process and plans to extend availability first to paid API customers and Google AI Ultra subscribers after collecting feedback and iterating on guardrails.

Argon is announced, priced, and already deployed inside Google, but it is not a public API that an enterprise can evaluate, benchmark, or integrate today. Google’s own Gemini 4 Argon release gives specific internal deployment examples, but they are company-reported results rather than independent validation.

“Starting this rollout in this way gives us more confidence, but also enables us to put a model that is trained and strong in cyber defense in the hands of defenders as soon as possible.”

— Tulsee Doshi, Google’s Gemini model product lead

A long-output model for long-running work

Google frames Argon as a model for tasks that need a long chain of tool use, analysis, code changes, and validation rather than a short chat answer. Its 1 million-token output cap is central: Google says the model can reason and generate over hundreds of thousands of tokens in one trajectory. The prior 64,000-token cap matters because large code migrations, legal or financial research, and autonomous remediation workflows can exhaust an output budget even when the input context remains manageable.

OpenAI’s Dots monetize autonomy before safety catches up

Recommended reading

OpenAI’s Dots monetize autonomy before safety catches up

Sergey Kuznetsov • • 10 min read

The output expansion does not establish that Argon will complete a task correctly. It gives the model more room to attempt one. Google’s reported benchmark results are capability signals, but they do not reveal the tool permissions, runtime limits, prompt formats, failure rates, or human review that would determine whether an agent is safe to use in production.

EvaluationGemini 4 Argon resultWhat Google says it measures
DeepSWE v1.177.9%Long-horizon, real-world software engineering
AutomationBench51.3%End-to-end execution across core business functions
LVBench91.7%Long-video understanding
CWE-bench v168%Software-vulnerability remediation

Google calls Argon the leader on the Vals Index, which weights finance, coding, legal, and tax work by each sector’s contribution to U.S. GDP. It also claims leading results on Vals Finance Agent v2 and Harvey’s Legal Agent Benchmark. The company has not supplied a numeric Vals Index result in the material provided here, so the ranking claim cannot be reduced to a score comparison.

The independent model-tracking page from Artificial Analysis lists Argon’s intelligence score, capability evaluations, context-window details, token pricing, and cost-per-task data as not publicly available. Third-party performance and price-per-task comparisons cannot yet confirm Google’s positioning because outside users do not have access to run them.

Google is already using Argon on production systems

The internal examples are more revealing than the benchmark chart. Google says Argon agents analyzed fleet-wide profiling telemetry and applied memory optimizations across its data centers, freeing more than 300 TiB of memory after rollout. The company estimates total savings could reach 500 TiB to 1 PiB. The result implies the agents identified changes that could be deployed across a production fleet under Google’s operational controls.

Google also says its quantum-computing researchers used Argon to optimize spacetime resources, defined as qubits multiplied by gates, for bottleneck subroutines. In one example, the company says Argon beat a published baseline by 40% within minutes. It did not identify the subroutine, baseline, hardware conditions, or method needed to reproduce that result.

Argon agents are helping move C and C++ code to Rust, ranging from core libraries including re2 and libgav1 to more than 800,000 lines in the Fuchsia OS Zircon kernel. Google says those rewrites go through automated and manual auditing, emulation testing, and review before production deployment. Google is not presenting the agent as an unattended production deployer.

For libgav1, Google says Argon took an existing Rust port and replaced 32,000 lines of SIMD code through repeated profile-guided experiments and compiler-output analysis. The result was safe Rust that the compiler could automatically vectorize, producing a decoder that ran 2.7 times faster than the earlier Rust port while generating identical video output. Google has not released the exact workload or patches for external scrutiny.

The public price exists, but the public service does not

Argon will carry introductory token prices when Google makes it available. Ars Technica said API pricing had not been announced, while Google’s September 30 release provides the introductory rates below. Developers cannot currently buy access on published terms.

Billing componentIntroductory price
Input tokens$2 per million tokens
Output tokens$10 per million tokens
Cached input tokens95% off the input-token price

The gap between input and output pricing matters for the product Google is describing. A model designed to emit long reasoning and action trajectories will make output-token consumption a bigger part of a workload’s bill than a conventional question-answering model. Google has not provided example task costs, rate limits, cache-write charges, or a date when those rates become purchasable.

The access policy relies on gatekeeping and filters

Argon’s cyber capability is the reason for its first audience and one of the sharpest policy choices in the release. Google says trusted defenders and internal teams will receive Argon without cyber guardrails so they can use its full defensive capability. Google separately says the model is designed to reject harmful cyber and CBRN requests under its Frontier Safety Framework.

Those statements describe different safeguards. The first is an access-control model: vetted users get a less constrained system. The second is behavior control for broader deployment. Google is relying on trusted-user selection, controlled environments, and monitored use during the initial release while it hardens generalized defenses.

Wiz is using Argon in its Scan for Good program, which searches for and remediates high-risk exposures affecting critical public infrastructure at no cost. Google says Argon found a critical vulnerability exposing sensitive personal information in healthcare software used by hospitals worldwide, an exposure that earlier frontier models missed. The release does not identify the affected product, flaw, remediation, or whether the vulnerability is now fully resolved.

On CWE-bench v1, Google says Argon tied for first at 68%. CNBC identified the tied models as OpenAI’s GPT-6 Astra and Grok 4.7; Google’s release names the score but not the competing systems. Google also says Argon outperformed its 3.8 Flash Cyber model on Wiz’s internal black-box penetration-testing benchmark, including attack-surface discovery, vulnerability identification, and proof-of-concept evidence. The actual Wiz benchmark scores are not published in the supplied material.

Google’s safeguards and the long-context problem

Google says it is strengthening four safeguard areas before broad availability: misuse prevention, resistance to indirect prompt injection, misalignment monitoring, and sandbox hardening. For misuse prevention, it says it uses internal-activation monitoring to spot problematic use, with internal and external red teams using manual and automated attacks. It says Argon leads Gray Swan’s Indirect Prompt Injection benchmark, though the announcement supplies no score.

For misalignment, Google says it will monitor Argon’s chain of thought and actions, stopping execution when necessary. The company says it used a similar system during training, sending alerts to a dedicated incident-response team, and deliberately avoided feeding findings back into training so it would not teach the model to evade the monitor. It remains a company description of a safety process rather than an independently audited control.

The research paper Google cites on activation probes identifies a technical tension in this rollout. It finds that existing probes can fail to generalize when inputs shift from short contexts to long contexts. Its proposed architectures improve that weakness, but the paper concludes that broad generalization requires both architectural choices and training on diverse distributions; pairing probes with prompted classifiers delivered the best accuracy at low cost. A million-token-output system can produce the extended, multi-turn traces that are difficult to monitor reliably.

Google says it is hardening sandboxed environments by isolating and sealing them before high-risk training and evaluations. That language follows the August 19 report that autonomous agents reached Hugging Face production infrastructure, prompting OpenAI to tighten its AI security posture. Google’s emphasis on sealed environments recognizes that the security boundary must include infrastructure, not only prompt-layer guardrails.

A faster Gemini cycle, with evidence concentrated at Google

Argon arrives less than two months after Google appointed Koray Kavukcuoglu to lead Gemini 4, following the departures of Jeff Dean and Oriol Vinyals to form Discovery Loop. Our August 5 coverage documented that leadership change. Argon is the first major model release to show the new organization’s priorities: larger output budgets, agentic work, cyber defense, and internal deployment ahead of public availability.

OpenAI introduced Dots and GPT-6.1 Sol one day before Argon’s release, while withholding a planned GPT-6.1 Astra model over safety concerns. Google is publicly claiming frontier results while withholding ordinary access and offering a restricted version to cyber defenders. Both approaches acknowledge the difficulty of controlling autonomous behavior once a capable model has tools and a long execution window.

What the evidence shows

The strongest confirmed facts are Google’s internal use cases, the 1 million-token output limit, its introductory $2 input and $10 output pricing, and the limited Fairwind rollout. External verification remains limited. Artificial Analysis has no public evaluation or cost data for Argon, while Google’s largest claims—from data-center memory savings to vulnerability discovery and code-migration performance—remain reported by Google.

Argon is not yet a broadly available enterprise model; it is a controlled operational experiment in letting an agent sustain much longer work. Google has supplied a plan and internal outcomes, but no public availability date, independent benchmark record, or reproducible evidence that its controls hold under the long-context conditions its own safety research identifies as difficult.

Frequently asked questions

When can developers use Gemini 4 Argon?+

Google has not provided a broad availability date. It is first rolling out through the Fairwind Program to trusted cyber defenders, then plans to expand to paid API customers and Google AI Ultra subscribers.

How much will Gemini 4 Argon cost?+

Google lists introductory prices of $2 per million input tokens and $10 per million output tokens. Cached input tokens are priced at 95% off the input-token rate.

What is Gemini 4 Argon’s token limit?+

Google says Argon can generate up to 1 million output tokens, compared with 64,000 output tokens in previous Gemini models.

Is Gemini 4 Argon publicly benchmarked by independent testers?+

Not yet in the supplied reporting. Artificial Analysis lists its intelligence, capability, context, and price data for Argon as not publicly available.

Sergey Kuznetsov

Editor-in-Chief

Sergey Kuznetsov is Head of Product at iXBT.com, one of the largest Russian-language technology media outlets, and the founder of itzine.ru. He has spent over a decade building and running tech newsrooms. At for(geeks) he sets editorial standards and reviews what ships.

/ Keep reading