Dev

Gemini 4 Argon's defender build is not the one you will call

Google is splitting access to its new frontier model by cyber safeguards, the way it did with 3.8 Flash and Anthropic did with Mythos. The developer build is still being tuned, so the benchmark table describes a model no developer will get in that form.

Google announced Gemini 4 Argon on September 30 with a benchmark table, a price and an output limit of a million tokens. Who can call it is another matter. Argon is rolling out to "trusted cyber defenders" through Google's Fairwind Program, and developers get it later, on a date nobody has named.

Read as a queue, that means developers get the same model a few weeks late. Google's own post says otherwise, and that changes what parts of the table are worth to you.

Can a developer call Gemini 4 Argon today?

No. The announcement says access will widen "as soon as possible", starting with paid API customers and Google AI Ultra subscribers. As of October 1 the Gemini API changelog has no Argon entry and no model ID. There's a second gate with no date on it either: Google says it is in the U.S. government's voluntary pre-release access process while access expands.

The price is already public: $2 per million input tokens and $10 per million output, cached input at 95% off. A footnote adds that $4 and $20 apply "after the introductory period expires", and the post never says when that is. So you can price a model you cannot run, at a rate with no end date. The standard rate is Opus 5.5's list price; the introductory one is Sonnet 5.5's.

Is "defenders first" just a delay?

It's a split by safeguards, and Google says so itself:

text
For trusted defenders and our own internal teams at Google, we'll be releasing Argon without cyber guardrails so they can leverage its full frontier-level cybersecurity defense capabilities.

Everyone else gets Argon after Google iterates "on guardrails". Google has run this split before. On September 2 it shipped Gemini 3.8 Flash and 3.8 Flash Cyber on the same day, "powered by the same foundational intelligence", with the Cyber variant carrying "a more permissive set of mitigations" and going only to Fairwind partners. With Argon, only the permissive side has shipped.

Anthropic took the same route this year. Project Glasswing gave Claude Mythos Preview to a closed group of partners on April 7, and Anthropic said it did not plan to make Preview generally available. On June 9 it launched Fable 5 and Mythos 5 together: one underlying model, safeguards lifted on Mythos, Fable available everywhere. Fable hands flagged cyber, bio and distillation requests to Opus 4.8, which Anthropic says happens in fewer than 5% of sessions. The Mythos 5.1 page now calls it "Claude Fable 5.1 for Project Glasswing participants", invite only.

So when Argon reaches the API, it will be the guarded version, and that version isn't finished yet.

What does the benchmark table tell a developer?

Google's results table has Argon at 77.9% on DeepSWE v1.1, 68.9% on the Vals Index, 51.3% on AutomationBench and 68% on CWE-bench v1, where it ties GPT-6 Astra. On Gray Swan's indirect prompt injection test the Fairwind page reports a 0.7% attack success rate at k=15. Several of these come from outside operators: Vals AI, Zapier's leaderboard, the CWE-bench leaderboard, Gray Swan. None of them can be rerun by a developer without access.

The methodology note is worth the five pages. Two computer-use rows ran "with safety filters enabled", flagged responses returned as empty strings. For the cyber rows it says nothing about guardrails either way. And Argon's DeepSWE score is self-computed with a mini-swe-agent harness, while the Fable 5.1 and Opus 5.5 numbers beside it come from their own system cards. The Vals lead over Opus 5.5 is 68.9 to 67.0.

Whichever build ran the table, it wasn't the one you'll call, because that one is still being tuned. For DeepSWE the difference is probably small. For CWE-bench and the prompt-injection result it is the whole question: a guarded model that declines part of a security task, or hands it to something weaker the way Fable does, scores differently from one that doesn't.

What should a roadmap that says "wait for Argon" do?

Treat it as a dependency with unknown terms. My rule for a model that won't hold still was to pin the version and fix the metrics in advance. A gated launch is the case where you can't pin anything yet. The model ID, the guardrails and the length of the price window all arrive after you've planned around them.

Don't pause a migration you'd otherwise make; the models you can call today are the only ones you can measure. Build the Argon eval now. The canary from the Opus 5.5 post is the skeleton: your real request, your parser, a loud failure. Add the cheap numbers you already watch (I use output length and structure count), and for security-adjacent tasks like dependency audits or auth code review, count what comes back refused or degraded. Then Argon's first day is a model ID swap.

The table won't settle role fit either. When I moved my scheduled agents to Opus 5, the more agentic model fit the reviewer's seat better than the worker's, and longer output was part of what hurt the bounded workers. A model built to write a million tokens in one response has a disposition too. No vendor table tells you which seat it fits.

There is no model ID to put in a config yet. The eval doesn't need one.

Discussion

No comment section here — all discussions happen on X.

Max Nardit

Max Nardit

@mnardit

More articles

Three records of one agent incident

OpenAI's logs, a URL scanner's public archive and a statistics bureau each tell it differently. Whoever keeps the fullest account is also the one grading it.

AI search visibility measurement framework for technical teams

One panel gives you a snapshot. A report compares two periods, and that is where most AI visibility changes turn out to be artifacts of the panel, the sample size or an unlogged model swap. The design that makes a trend checkable.

Sonnet 5.5 keeps its reasoning for one model and one account

Switch to a model or an API account that can't read the earlier thinking, and the request still succeeds. The reasoning those turns built up is dropped before the model reads it, with no error and nothing on the bill. You only find out if you asked the API to report it.