Home AI Gemini 4 Argon Is Google’s New Frontier Model, but Most People Can’t...

Gemini 4 Argon Is Google’s New Frontier Model, but Most People Can’t Use It Yet

Argon tops many of Google's benchmarks and trails on others. The $2/$10 price is introductory, and access starts with cyber defenders.

0
Gemini 4 Argon logo next to a large numeral 4 on a blue gradient background
Image: Google

Google has announced Gemini 4 Argon, its new frontier AI model and the first in the Gemini 4 family. However, most people cannot use it yet. Gemini 4 Argon is rolling out first to a small group of vetted cyber defenders, and Google has not given a date for wider access.

Google says the model is built to sustain deep reasoning across long, complex jobs. It targets software engineering, legal and finance work, and cybersecurity defense.

Gemini 4 Argon pricing has an asterisk

Argon launches at $2 per million input tokens and $10 per million output tokens. That is an introductory rate, though. Google’s footnote says the price doubles to $4 and $20 once the introductory period ends, and it has not said when that happens. Cached input tokens get a 95% discount.

The model also gets a much larger output limit. Google raised it to 1 million tokens, up from 64,000, so Argon can work through a hard problem in a single long run.

Where Argon leads, and where it does not

Google’s benchmark table shows Argon ahead in most rows. For example, it scores 77.9% on DeepSWE v1.1, a test of long software engineering tasks. GPT-6 Astra scores 74.1%, and Claude Opus 5.5 scores 74.2%. Argon also tops AutomationBench at 51.3% and leads the Vals Index for finance, legal, and tax work. In addition, it scores 91.7% on LVBench, a test of long video understanding.

Google’s own comparison table. Shaded cells mark the top score in each row. Image: Google

Still, the same table shows clear losses. Argon scores 55.0% on FrontierSWE v2, well behind GPT-6 Astra at 65.5%. On Terminal-bench 4.0, it trails Claude Opus 5.5, which scores 66.4% to Argon’s 57.4%. GPT-6 Astra also wins on Terminal-Bench Science and OSWorld 2.0.

Independent testing is more measured, too. Artificial Analysis gives Argon a score of 53 on its Intelligence Index, level with GPT-6 Astra. At the discounted price, it costs about 60% as much per task as Astra. After the discount ends, however, the firm estimates Argon will cost roughly 1.2 times as much as Astra.

One result stands out. Artificial Analysis measured a 15% hallucination rate for Argon, compared with 51% for GPT-6 Astra. In practice, that means Argon is more likely to admit it does not know an answer. Its accuracy on the same test was lower than Astra’s, though, at 50% versus 63%.

A cybersecurity model with fewer guardrails for defenders

Google trained Argon heavily for cyber defense. “Argon can autonomously find, validate, and patch critical software vulnerabilities,” the company writes. For trusted defenders and its own teams, Google will release the model without cyber guardrails.

Argon ties Grok 4.7 and GPT-6 Astra at 68% on CWE-bench v1. Image: Google

On CWE-bench v1, which measures how well models fix security flaws, Argon ties for first at 68%. Google also says security firm Wiz used the model to find a critical vulnerability in healthcare software that hospitals use worldwide. According to Google, earlier frontier models had missed it.

That power is why the rollout is slow. Google says it is still strengthening safeguards against misuse, prompt injection, and misalignment. It will monitor Argon’s chain of thought and actions, and stop execution when needed. On Gray Swan’s prompt injection benchmark, Google reports the lowest attack success rate of any model it compared. The company is also taking part in the U.S. government’s voluntary pre-release testing process.

Who gets it, and when

For now, access runs through Google’s Fairwind Program for trusted cyber defenders. After that, Google plans to open Argon to developers, enterprises, and consumers, starting with paid API customers and Google AI Ultra subscribers. The company only says that will happen “as soon as possible.”

Meanwhile, Google is already using the model itself. It says Argon agents freed more than 300 TiB of memory across its data centers and are helping move C and C++ code to Rust. Until the model is widely available, though, most of those claims rest on Google’s own testing.

NO COMMENTS

Exit mobile version