On September 30, Google launched Gemini 4 Argon, its latest flagship AI model designed for programming, corporate tasks, and cybersecurity, focusing on executing long, multi-step tasks.

A comparison of Gemini 4 Argon's capabilities across various benchmarks against other leading AI models. Source: Google Blog.

Currently, Gemini 4 Argon is undergoing testing by select security specialists through the Fairwind program. The company is also participating in a voluntary pre-screening initiative by the U.S. government for advanced models.

Following this phase, the model will first be accessible to paying API users and subscribers of Google AI Ultra, although the company has not disclosed the timeline for a public release.

Token Output Increased to 1 Million

Google has raised the maximum output limit from 64,000 to 1 million tokens. According to the developers, this enhancement allows the model to better handle complex and time-consuming tasks within a single sequence and generate hundreds of thousands of tokens continuously.

In the DeepSWE v1.1 benchmark, which assesses prolonged engagement with real software projects, Argon achieved a score of 77.9%.

Gemini 4 Argon's performance in DeepSWE v1.1. Source: Google Blog.

However, performance varies by task. In the FrontierSWE v2 benchmark, the model scored 55%, compared to 65.5% for GPT-6 Astra and 62.3% for Claude Opus 5.5.

Results of leading AI models in the FrontierSWE v2 benchmark. Source: FrontierSWE.

In corporate tasks, Argon scored 65.4% in Vals Finance Agent v2 and 51.3% in AutomationBench from Zapier. It also achieved a score of 91.7% in the LVBench test for understanding lengthy videos.

The initial pricing for the API will be $2 per 1 million input tokens and $10 per 1 million output tokens. Caching input will be significantly cheaper at $0.10 per 1 million tokens.

After the introductory period, prices will rise to $4 and $20, respectively, although Google has not specified when this will occur.

Internal Use of Argon at Google

According to statements, the flagship model is already being utilized by thousands of Google employees. One group of agents analyzed data center operations, identifying ways to free up over 300 TiB of RAM.

The potential savings from broader implementation are estimated by the company to be between 500 TiB and 1 PiB.

Other agents are assisting in porting code from C/C++ to Rust, with project sizes ranging from libraries containing tens of thousands of lines to the Zircon kernel of the Fuchsia operating system, which exceeds 800,000 lines.

While working with the libgav1 video decoder, Argon processed around 32,000 lines of SIMD code. According to Google, the final Rust version operates 2.7 times faster than the previous port while delivering the same results.

The company also claims that in one experiment, the model improved an existing quantum algorithm optimization result by 40% in just a few minutes.

Focus on Cybersecurity

Argon has been trained to autonomously identify, verify, and rectify software vulnerabilities.

In the CWE-bench v1 benchmark, the model received a pass@1 score of 68%. This is on par with the current performance of GPT-6 Astra and Grok 4.7.

Performance of leading AI models in the CWE-bench v1 benchmark. Source: CWE-bench.

The company Wiz is already testing Argon in its Scan for Good program. Google noted that the model discovered a critical vulnerability in medical software used by hospitals globally, which could have exposed sensitive personal data; other advanced models failed to detect this issue.

Prior to a wider rollout, Google plans to enhance protections against misuse and attacks through hidden instructions.

Additionally, Google has implemented a system that monitors the model's operations and can halt tasks if its actions exceed user intentions.

In September, Google had also introduced Gemini 3.8 Flash for extended programming tasks and autonomous agents, along with a specialized version, Gemini 3.8 Flash Cyber, aimed at identifying and rectifying vulnerabilities.

Later, it was reported that during tests of Gemini's cybersecurity capabilities, the model accessed the systems of three real companies instead of a simulated target. Google stated that the model autonomously stopped after detecting an error and did not cause any harm.

Follow ForkLog on social media

Telegram (main channel) Facebook X