AI News

Gemini 4 Argon: Specs, Benchmarks, Pricing and Model Comparisons

Google DeepMind announced Gemini 4 Argon on September 30, 2026, with a clear focus on sustained professional work: software engineering, enterprise research, legal and financial workflows, and defensive cybersecurity. Its most striking technical disclosure is a one-million-token output limit. Its most persuasive benchmark results cluster around knowledge work, long-context reasoning and multimodal understanding. Its initial release is restricted to trusted testers.

That combination makes Argon a significant release to watch. It also makes the details essential. A million output tokens are different from a million tokens of input context. Introductory prices are different from permanent prices. A model that leads DeepSWE can still trail competing models on another software-engineering evaluation.

This guide examines the announced specifications, all 19 rows in Google’s main comparison table, the evaluation conditions, API economics, internal deployment examples and practical differences against GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5. Google has published results; Kingy.ai has not independently run Argon. Claims about internal deployments and safety performance are attributed to Google throughout. Sources were checked on September 30, 2026. Google’s launch announcement and official performance table are the primary references.

Gemini 4 Argon is announced, with a phased release

The linked Google DeepMind announcement on X is genuine. Google identifies Argon as its new frontier model, and the accompanying official blog describes access beginning with trusted cyber defenders through the Fairwind Program.

Google says broader access will follow for developers, enterprises and consumers, starting with paid API customers and Google AI Ultra subscribers. It does not give a firm public rollout date in the launch announcement. A published price therefore establishes the intended commercial rate; it does not establish that any developer can call the model today.

Google also says it is participating in the U.S. government’s voluntary process for pre-release model access and using feedback from early testers to refine safeguards. The practical distinction is between an announced model, access for a selected cohort, and general availability. Argon has reached the first two stages at launch. Readers should verify their own account’s eligibility when the broader release arrives. Source: Google’s rollout section.

The specifications Google has actually disclosed

Gemini 4 Argon launch disclosures, checked September 30, 2026
Item Verified disclosure Practical interpretation
Model name Gemini 4 Argon Google’s official announced name; do not substitute an assumed API identifier.
Output token limit 1 million tokens, up from Google’s stated previous 64K limit More room for long generation and reasoning trajectories; not a guarantee of useful output at that length.
Input context limit No definitive commercial input limit stated in the launch announcement Google reports long-context evaluation results through 1M tokens, but benchmark coverage does not settle the eventual API limit.
Primary workloads Software engineering, enterprise knowledge work and cybersecurity defense Evaluate whole workflows, including tools and validation, rather than chat fluency alone.
Multimodal evidence Published chart-understanding and long-video results Understanding capability does not establish image, video or audio generation endpoints.
Initial access Trusted cyber defenders through Fairwind; Google internal teams Public availability remains phased.
Broader access plan Paid API customers and Google AI Ultra subscribers first No specific public launch date or subscriber usage allowance is given.
Introductory API rate $2 per million input tokens; $10 per million output tokens Published launch pricing, with a later increase explicitly disclosed.
Cached input 95% discount against input token price $0.10 per million cached input tokens at the introductory rate, calculated from the stated discount.
After introductory period $4 input; $20 output per million tokens Both headline rates double; the announcement does not specify the introductory end date.
Undisclosed deployment details Parameter count, architecture details, training compute, public model ID, latency and throughput are not specified in the launch article Leave these fields open until documentation supplies them.

The table combines direct disclosures with explicitly labeled arithmetic and interpretation. It deliberately leaves uncertain fields uncertain. Google has not supplied enough information in this announcement to write a conventional hardware-style spec sheet with parameter count, model size, tokens per second and every supported API modality. Source: official launch announcement.

The same care applies to ecosystem features. DeepMind’s Gemini page also links to Gemini Omni, Nano Banana, Gemini Audio and Gemini Robotics. Those are separate product and model surfaces. Their presence on the page does not mean Argon includes every generation or robotics feature they advertise.

What a one-million-token output limit changes

Most discussion of large token limits concerns input: how much source code, documentation, video or conversation a model can read. Google specifically calls Argon’s new limit an output limit. The model has more room to produce a long reasoning and generation trajectory while working through a difficult problem. Google connects that room to deeper reasoning across longer tasks. Source: the output-limit announcement.

The distinction matters for coding agents. A difficult migration may require inspecting a repository, proposing changes, running tests, examining compiler output, revising the approach and checking another set of failures. A larger generation budget can reduce the pressure to terminate a trajectory early. The potential benefit is more sustained work before the model reaches its generation ceiling.

It does not guarantee that an agent remembers every earlier detail, chooses the right tool, produces a correct patch or finishes faster. Those depend on the model, its input context, the surrounding agent software, the task and the validation process. A high ceiling can also permit an expensive unsuccessful attempt.

Google’s comparison from 64K to 1M represents roughly a sixteenfold increase. The exact ratio depends on whether the abbreviated older limit means 64,000 or 65,536 tokens; the release uses the shorthand rather than resolving that distinction. Either way, the change is substantial.

A million-token ceiling should not become the default budget for every request. A short classification, a retrieval-backed answer and a multi-hour repository task need different limits. An application should cap generation to the workload, record actual token consumption and stop attempts that keep spending without making measurable progress. These are implementation recommendations derived from the announced limit, not measurements of Argon’s behavior.

There is also an unresolved billing detail: the launch blog does not fully document how reasoning tokens, visible output, tool activity and the maximum generation setting will map onto the public API. The cost examples below use billable output tokens as an explicit assumption. They should be updated against the API documentation when broader access opens.

The complete published comparison against Astra and Claude

Google’s main table compares Argon with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5. The rows cover several distinct capabilities. Higher is better for the displayed scores, but the percentages are not interchangeable: some measure successful tasks, others measure accuracy or F1, and OSWorld uses a partial score.

The following tables reproduce the reported numerical results. They are Google’s launch comparison, not an independent Kingy.ai benchmark. A dash indicates a result Google did not include. Exact evaluation versions and conditions matter; the methodology section below explains the largest comparability limits. Source: Google DeepMind’s performance table.

Enterprise knowledge work

Benchmark Argon GPT-6 Astra Fable 5.1 Opus 5.5
Vals Index 68.9% 63.1% 65.8% 67.0%
AutomationBench 51.3% 41.4% 31.4% 42.5%
Vals Finance Agent v2 65.4% 53.5% 58.9% 58.6%
Harvey’s Legal Agent Benchmark 19.6% 5.4% 6.7% 3.8%

Argon leads all four displayed knowledge-work rows. The Vals Index margin is relatively narrow: 1.9 percentage points above Opus 5.5. The AutomationBench margin is larger, at 8.8 points over Opus, the next highest displayed result. On Vals Finance Agent v2, Argon leads Fable by 6.5 points and Opus by 6.8 points.

Google describes the Vals Index as a GDP-weighted measure spanning finance, coding, legal and tax tasks. That weighting makes it different from an equal-weight average of academic tests. It gives the index an economic framing, but it does not mean 68.9% of a company’s work can be automated.

AutomationBench addresses end-to-end execution across business functions, according to Google’s description of Zapier’s evaluation. A 51.3% result is a lead in this comparison, while also showing substantial room for failure under the test conditions. Enterprise buyers should care about both statements. A system used for business workflows needs reliable recovery and validation when an attempt fails.

The legal result deserves particular care. Argon’s 19.6% is roughly 2.9 times Fable’s 6.7%, but the absolute score remains 19.6%. A large relative improvement from a low base can sound more conclusive than it is. This benchmark does not establish legal accuracy for every jurisdiction or permission to publish filings without professional review. The scores measure the benchmark’s tasks and grading rules.

Software engineering and terminal work

Benchmark Argon GPT-6 Astra Fable 5.1 Opus 5.5
DeepSWE v1.1 77.9% 74.1% 67.4% 74.2%
FrontierSWE v2 55.0% 65.5% 56.3% 62.3%
Vibe Code Bench 91.9% 89.6% 90.3% 90.3%
Terminal-bench 4.0 57.4% 58.2% 57.9% 66.4%

The coding evidence is mixed. Argon leads DeepSWE v1.1 by 3.7 percentage points over Opus and 3.8 points over Astra. It also leads Vibe Code Bench by 1.6 points over both Claude models. Those are useful results for a model positioned around sustained engineering work.

On FrontierSWE v2, however, Astra’s 65.5% is 10.5 points above Argon’s 55.0%. Opus is 7.3 points ahead, and Fable is 1.3 points ahead. On Terminal-bench 4.0, Opus scores 66.4% against Argon’s 57.4%, a nine-point lead. Astra and Fable also have slightly higher displayed terminal scores.

The difference between DeepSWE and FrontierSWE should shape buying decisions. Software-engineering evaluations can sample different repositories, task lengths, languages and required tool behavior. A lead on one suite establishes a lead under that suite’s conditions. It does not erase a deficit on another suite.

For a team migrating a large codebase, DeepSWE and Google’s internal migration examples make Argon a plausible candidate for a controlled trial. For a team whose workload resembles FrontierSWE or terminal-heavy tasks, the displayed Astra and Opus results supply a strong reason to keep those models in the comparison. Match the trial to the actual work rather than picking the model from one favorable row.

Machine-learning engineering, science and mathematics

Benchmark Argon GPT-6 Astra Fable 5.1 Opus 5.5
PostTrainBench 45.3% 44.3% 40.2% 49.3%
Terminal-Bench Science 0.1 57.6% 68.1% 52.6% 63.3%
LABBench 2 88.8% 85.4% 68.6% 73.1%
RiemannBench 76.0% 72.0% 65.6% 69.6%

Argon leads LABBench 2 by 3.4 points over Astra and RiemannBench by four points over Astra. The two wins support Google’s claim of broad reasoning strength. The other two rows qualify it: Opus leads PostTrainBench by four points, and Astra leads Terminal-Bench Science 0.1 by 10.5 points.

Post-training engineering, scientific terminal tasks, laboratory-related reasoning and mathematical reasoning are different activities. The pattern suggests Argon can be competitive across them while still leaving model choice dependent on the task. The methodology also uses different verifier timeouts on Terminal-Bench Science, which makes a simple score-only comparison incomplete.

Long-context reasoning

GraphWalks condition Argon GPT-6 Astra Fable 5.1 Opus 5.5
Up to 128K; breadth-first search, F1 99.7% 98.7% 91.4% 90.6%
256K to 1M; breadth-first search, F1 84.2% 71.8% 65.0% 66.8%

The longer-context row is one of Argon’s clearest displayed advantages. Its 84.2% F1 is 12.4 percentage points above Astra, 17.4 above Opus and 19.2 above Fable. At up to 128K, the gap to Astra is only one point.

That pattern matters because large nominal context windows can conceal a decline in reasoning quality as input gets longer. GraphWalks asks for structured reasoning over supplied graph information. Argon’s result is evidence for that particular long-context reasoning task. It is more informative than a context-size claim alone, but it still does not guarantee perfect retrieval or reasoning over every million-token repository or document collection.

Keep the two million-token discussions separate: the GraphWalks row measures performance over long inputs, while the launch announcement’s 1M specification concerns output. Google has disclosed both a long-context evaluation result and a generation ceiling. It has not turned those two disclosures into a complete public API specification.

Computer use, multimodal understanding and cybersecurity

Benchmark and condition Argon GPT-6 Astra Fable 5.1 Opus 5.5
Agent’s Last Exam; pass rate 39.5% 34.2% — 38.2%
OSWorld-2.0; offline subset, partial score 69.2% 72.6% — —
Chartography 71.6% 71.0% 46.2% 66.3%
LVBench 91.7% 87.5% 79.7% 83.7%
CWE-bench v1 68.0% 68.0% 58.0% 67.0%

Argon leads the displayed Agent’s Last Exam results, but its 39.5% pass rate remains below half. On the OSWorld-2.0 offline subset, Astra leads by 3.4 points. The omitted Claude results must remain omitted: Google’s methodology says comparable offline-only results were unavailable, so their absence is not a zero score.

Chartography shows a narrow 0.6-point lead over Astra. LVBench shows a larger 4.2-point lead over Astra and an eight-point lead over Opus. Google describes LVBench as a long-video-understanding test. Different frame budgets across models affect how much visual evidence each model receives, so the score needs that qualification.

On CWE-bench v1, Argon ties Astra at 68.0%, one point above Opus. This is a vulnerability-remediation result. It should not be converted into a claim that Argon discovers 68% of every organization’s vulnerabilities or offers comprehensive security coverage.

Across the 19 displayed rows, Argon has the highest numerical result in 13, ties for highest in one and trails another model in five. That is a count of this particular table, calculated by Kingy.ai. It is not a weighted overall score: the table includes two GraphWalks conditions, incomplete rival coverage and several different grading measures. Averaging the percentages would imply a common scale they do not have.

How the evaluation conditions change the interpretation

Google’s five-page evaluation methodology is essential reading alongside the headline table. In general, Argon runs at the highest thinking settings, with single-attempt evaluation rather than majority voting or parallel test-time compute. Smaller benchmarks use repeated trials to reduce variance. Some rows are Google’s own experiments; others come from evaluator leaderboards or rival providers’ reports.

That means the launch table is a compilation of documented results, with benchmark-specific methods. It is not a single laboratory running every model through every test under one perfectly identical configuration. The following qualifications materially affect what the numbers can support.

  • DeepSWE uses mixed sources. Google computed Argon’s result with a mini-swe agent harness. Rival numbers come from a public leaderboard and model reports. Reasoning levels and the surrounding agent software are part of the comparison.
  • PostTrainBench has a useful shared setup. Google computed all four models’ v1.1 results with OpenCode, a ten-hour budget and a single NVIDIA H100. This is more controlled than mixing unrelated reports, but it remains a provider-run experiment.
  • LABBench 2 is tool-assisted. Models can use a Linux terminal, bioinformatics software, Python, R and internet access. Its scores represent work with that environment, rather than an unaided science quiz.
  • Terminal-Bench Science has a timeout difference. Argon uses a six-times verifier timeout to address verification timeout issues. The competing results come from the official leaderboard. The altered timeout belongs beside the comparison.
  • GraphWalks covers specific subsets. The shorter condition contains 650 items at up to 128K tokens; the longer condition contains 200 problems between 256K and 1M. F1 measures the requested graph reasoning, not complete correctness on arbitrary long documents.
  • OSWorld is an offline partial score. Argon’s result is the maximum over three runs, each with a single attempt, using screenshot observations at 1080p and a 500-step limit. Google’s table omits Claude because its available figures combine online and offline tasks. Treating these missing entries as losses would misrepresent the comparison.
  • LVBench uses unequal video sampling budgets. Google uses one frame per second for Gemini, 800 frames for Astra, 300 for Fable and 600 for Opus, reflecting API limits. This tests the available systems under those conditions; it does not give every model identical visual evidence.

Google’s methodology also contains a date inconsistency: its introductory text discusses September 2026 capabilities, while the final results page labels the figures as of October 2026. The announcement itself is dated September 30. We report the launch material as published and do not infer an additional month of testing.

The independent evaluator pages add useful precision. Vals Finance Agent v2 grades its primary score using weighted partial credit gated by dealbreaker requirements. Argon’s 65.4% therefore should not be read as 65.4% of financial tasks being completely correct. By contrast, Zapier’s AutomationBench uses strict completion requirements: every assertion for a task must pass. Identical-looking percentages can describe very different levels of success.

Finally, Google’s four-model selection leaves out some current competitors. Vals’ Vibe Code Bench page reports 92.39% for Claude Sonnet 5.5, above Argon’s 91.9% in Google’s table. On CWE-bench v1, Grok 4.7 also reaches 68% pass@1. Argon leads or ties many selected rows; the wider field prevents calling every one of those rows an exclusive overall win.

Gemini 4 Argon pricing, with the footnote included

Argon launches at $2 per million input tokens and $10 per million output tokens. Google discounts cached input tokens by 95%, which calculates to $0.10 per million cached input tokens at that introductory rate. After the introduction, the published input and output prices become $4 and $20. Google does not specify the introduction’s end date in this announcement. Source: launch pricing and footnote.

All prices below are U.S. dollars per million tokens. The rival entries are current standard base API prices, excluding premium service tiers, long-input premiums, batch discounts, cache writes, storage, tools and taxes. This compares unit prices; equal token counts are not guaranteed for the same work across providers.

Model or pricing phase Uncached input Cached input read Output
Gemini 4 Argon, introductory $2.00 $0.10, calculated $10.00
Gemini 4 Argon, after introduction $4.00 $0.20 if the same discount continues $20.00
GPT-6 Astra $10.00 $1.00 $50.00
GPT-6.1 Sol $2.00 $0.10 $10.00
Claude Fable 5.1 $10.00 $0.25 $50.00
Claude Opus 5.5 $4.00 $0.20 $20.00
Gemini 3.8 Flash, current introductory rate $0.75 $0.075 $3.75

Primary pricing references: GPT-6 Astra, GPT-6.1 Sol, Anthropic pricing and Gemini API pricing. Argon’s later cached-input figure is a conditional calculation, because the announcement does not separately enumerate that future cached rate.

At introductory rates, Argon’s uncached input and output are 80% cheaper per token than Astra and Fable’s $10/$50 base rates. After introduction, they are 60% cheaper. Argon’s introductory rates are half Opus 5.5’s, while its announced later rates match Opus.

GPT-6.1 Sol complicates a comparison focused only on the most expensive rivals. OpenAI’s current model documentation gives Sol the same $2/$10 rates and $0.10 cached-input read price as introductory Argon. Google’s launch benchmark table does not include Sol 6.1, so it cannot settle their relative quality. A practical value comparison should include it rather than assuming Argon is uniquely inexpensive among frontier candidates.

Existing Gemini users also have a lower-priced alternative in 3.8 Flash. Its current $0.75/$3.75 rates are below Argon’s introduction, with standard $1.50/$7.50 rates scheduled from January 1, 2027. That date belongs to Flash’s price schedule; it must not be imported as Argon’s introductory deadline.

Worked token-cost examples

A basic calculation is: uncached input tokens multiplied by the input rate, plus cached input tokens multiplied by the cached rate, plus billable output tokens multiplied by the output rate. Divide each token count by one million. Add any cache creation, storage, tool and infrastructure charges separately.

Hypothetical request Argon introduction Argon later rate What is assumed
100K uncached input + 20K output $0.40 $0.80 Input and output billed at the stated base rates.
100K cached input + 20K output $0.21 $0.42, conditional A valid cache hit; later 95% discount continues; no cache write/storage fees.
100K uncached input + 100K output $1.20 $2.40 Every generated token in this example is billable output.
1M output tokens alone $10.00 $20.00 Excludes input, tools, retries and all other fees.

These are arithmetic examples, not observed Argon invoices. At equal base-rate token counts, the first request costs $2.00 on Astra or Fable, $0.80 on Opus and $0.40 on Sol 6.1. Real tasks can use different token counts, take different numbers of attempts and produce different results.

The million-token example explains why a low unit price can coexist with a meaningful task bill. An agent that runs several long unsuccessful trajectories may cost more than a higher-priced model that finishes in one short attempt. Track cost per accepted result, elapsed time and the amount of human correction required.

Caching can strongly change repeated-input economics. A 100K-token prefix costs $0.20 as uncached Argon input at the introductory price, versus $0.01 as a qualifying cache read. Whether a real workflow gets those savings depends on cache eligibility, the hit rate and any creation or retention charges. Argon’s launch announcement does not provide a complete cache-fee schedule.

Long-input rules also differ across rivals. OpenAI’s Astra and Sol documentation applies doubled input and cached-input rates and a 1.5-times output rate to requests above 272K input tokens. Anthropic separately charges cache writes and offers service-specific pricing. These terms can change a long-context comparison enough that a four-number rate table is insufficient for budgeting.

Evaluator-reported task cost is more useful than token price alone

Zapier’s live AutomationBench results report Argon High at 51.29% and $1.70 per task at standard list prices, or $0.85 under promotional pricing. Argon Medium scores 50.08% at $1.54 list or $0.77 promotional. Opus 5.5 Max with default fallbacks scores 42.47% at $1.44 per task; Astra Max scores 41.4% at $1.73. These figures are specific to Zapier’s workload and its accounting.

That example shows both the commercial appeal and the future price change. At the promotional rate, Argon High costs less per task than these two rivals in the reported experiment. At the later list rate, its task cost is above Opus and close to Astra, while its completion score remains higher. None of those relationships follows from token prices alone.

The same leaderboard also reports Claude Sonnet 5.5 Max with default fallbacks at 44.75% and $1.14 per task. Its notes flag fallback use for some Claude configurations; Fable’s 31.4% includes Opus 5 fallback on roughly 40% of tasks. A result for a configured agent with fallback is a result for that system, rather than a clean measurement of a single model operating unaided.

CWE-bench v1 presents a different cost pattern. Argon, Astra and Grok 4.7 each score 68% pass@1, but their reported average rollout costs are $6.63, $2.85 and $2.75 respectively. Opus 5.5 scores 67% at $0.79. The evaluated harnesses differ: Antigravity for Argon, Codex for Astra, OpenCode for Grok and Claude Code for Opus. Close scores therefore coexist with materially different system costs.

CWE-bench’s reported Argon cost uses its own accounting assumptions, including a cached-token rate that differs from the introductory rate calculated from Google’s launch discount. We retain the evaluator’s published bill rather than retroactively recomputing it. Benchmark task cost is evidence about that experiment, not a universal price quote.

Google’s internal results: useful evidence with specific baselines

Google says thousands of employees have used Argon for specialized coding, research and writing. It also publishes several engineering examples. These are provider-reported deployment stories, rather than reproducible head-to-head trials against the three models in the benchmark table. Their value comes from the specificity of the task and baseline. Source: Google’s internal workflow examples.

In quantum computing, Google says Argon improved a published baseline for the spacetime resources of a subroutine by 40% in minutes. The resource measure combines qubits and gates. This is a claim about a particular optimization problem, not a 40% improvement in all quantum computing workloads.

For data-center memory, Google reports that a team of agents inspected fleet profiling telemetry and applied optimizations that freed more than 300 TiB once deployed. It estimates total potential savings of 500 TiB to 1 PiB. The distinction between already freed memory and projected total savings belongs in any retelling of the result.

Google also describes C/C++ to Rust migrations spanning core libraries such as re2 and libgav1, up to more than 800,000 lines for Fuchsia’s Zircon kernel. It says critical rewrites undergo automated and manual audits, emulation testing and review before production. The announcement does not say Argon independently replaced the entire kernel and deployed it without review.

The libgav1 example is especially concrete. Argon agents took an existing Rust port and replaced 32,000 lines of SIMD code through repeated profile-guided experiments and compiler inspection. Google reports an output-equivalent decoder that runs 2.7 times faster than that Rust port. The comparator is the earlier Rust port; Google says the result moves closer to optimized C++. Calling it 2.7 times faster than the C++ decoder would change the claim.

For engineering teams, these examples suggest what a useful Argon trial might look like: a defined repository, an explicit baseline, performance measurements, output-equivalence checks and mandatory review of the resulting changes. They also show the surrounding infrastructure required to turn model output into production improvement.

Cybersecurity capability, controlled access and safety evaluations

Cyber defense is central to Argon’s initial release. Google says the model can find, validate and patch software vulnerabilities. It describes work with Wiz’s Scan for Good initiative, including discovery of a critical exposure in healthcare software that earlier frontier models missed. The launch article does not provide enough public technical evidence to reproduce that specific discovery, so it remains an attributed example.

Google reports two additional internal security evaluations against Gemini 3.8 Flash Cyber. In source-code vulnerability discovery across 20 programming languages, Argon reaches 85.8% compared with 71.0%, a 14.8-point gain. In Wiz’s black-box penetration-test evaluation, it reaches 70.9% compared with 58.2%, a 12.7-point gain. The first measures recall of recent confirmed vulnerabilities in source code; the second examines web systems without source-code access. They address different tasks. Source: Google’s security evaluation chart and methodology.

Those internal figures should remain distinct from the public CWE-bench results. They do not establish that Argon discovers 85.8% of unknown zero-days, nor that its larger recall percentage can be compared directly with a 68% remediation score.

Google says trusted cyber defenders and its internal teams will receive Argon without cyber guardrails to support authorized defensive work. The broader release is described with safeguards against misuse. These are different access conditions; the restricted deployment should not be presented as the ordinary consumer experience.

The Fairwind Program page describes more than 650 partners globally, but the launch announcement gives Argon to a selected set. Program membership does not establish that every partner has immediate Argon access. The program includes due diligence and defensive-use restrictions.

Prompt-injection results

Google’s launch chart reports the following attack-success rates on Gray Swan’s Indirect Prompt Injection evaluation at 15 attack attempts. Lower is better. This is a measured attack-success metric, not a general percentage of safety.

Model Attack success at k=15
Gemini 4 Argon 0.7%
Claude Opus 5.5 1.0%
Claude Fable 5.1 1.0%
Gemini 3.8 Flash 5.5%
Gemini 3.8 Flash Cyber 6.0%
GPT-6 Astra 8.5%

Source: Google’s Gray Swan comparison chart. Argon’s 0.7% is the lowest displayed result in that chart. The margin over Opus and Fable is 0.3 percentage points. The chart does not supply enough uncertainty information to treat that small numerical difference as a settled statistical advantage.

Indirect prompt injection is especially relevant to agents that read web pages, repositories, messages or documents. An attacker can place instructions inside material the agent is meant to inspect. Google’s reported result is encouraging evidence under one adversarial test. It does not establish immunity across every tool, permission configuration or attack surface.

What Google says it is doing before broad release

Google groups its safeguards into four areas: refusing harmful requests while preserving legitimate research, resisting prompt injection, monitoring misalignment and hardening the environments in which powerful models operate. It describes internal and external red teaming, monitoring internal activations, and checking reasoning and actions for behavior that exceeds the user’s intentions. Source: Google’s safeguard disclosures.

The reasoning-monitoring claim needs a precise reading. Google says it monitors chain-of-thought and actions and can stop execution when necessary. Google says a similar system monitored training runs and sent alerts to a dedicated incident response team, with precautions against feeding monitoring findings back into training in a way that encourages evasion. This is a description of its mitigation approach, not proof that reasoning traces always reveal every harmful intention.

For deployments, model safeguards should sit alongside scoped permissions, isolated execution, audit logs and review of consequential actions. Argon’s longer task horizon increases the importance of the agent’s control system, because the model may take many steps before a person checks the final result. The launch’s restricted access and testing phase reflect the need to evaluate those systems together.

How Argon compares in a practical model selection

Argon’s strongest case in the published evidence is a workflow that combines sustained reasoning, knowledge work and large amounts of visual or textual evidence. Its leads on AutomationBench, Finance Agent v2, GraphWalks and LVBench make those sensible starting points for evaluation. Its introductory pricing adds a reason to test it once access is available.

Astra remains competitive where the task resembles FrontierSWE v2, Terminal-Bench Science or the OSWorld offline subset. Opus remains competitive for terminal work and post-training engineering. Fable has a higher displayed API rate than introductory Argon, but a purchasing decision should still test the actual task, reasoning configuration and required quality; the headline table does not measure every workload.

Broaden the candidate set when value matters. GPT-6.1 Sol matches Argon’s introductory unit prices and has public documentation. Claude Sonnet 5.5 appears on independent leaderboards and can beat models in Google’s selected comparison on particular tests. Gemini 3.8 Flash has lower current unit rates. Grok 4.7 shares Argon’s CWE-bench pass@1 score. None of these observations supplies a universal winner; each identifies a comparison the launch table leaves open.

Current public API specifications also give context to Argon’s output headline. Astra and Sol 6.1 document 1,050,000-token context windows and 128,000-token normal output limits. Fable 5.1 and Opus 5.5 document 1M context windows and 128K normal output; Opus additionally offers up to 300K output through a Message Batches beta. Argon’s advertised million-token output ceiling is much larger, but its public API contract and normal operating limits still need documentation. Astra specifications, Sol specifications, Fable specifications and Opus specifications.

A useful trial should contain representative tasks with acceptance criteria established before any model runs. For code, measure test success, regression risk, patch quality and reviewer time. For document research, check citations, omitted evidence and material errors. For automation, check whether the final application state meets every requirement. For multimodal work, record the frame or image budgets as well as the answer score.

Run the same tools and permissions where possible, record reasoning settings, and report both cost and elapsed time. Save failed attempts as well as successes. A model that produces a correct result occasionally at low token cost can still be expensive after retries and manual repair. A model with a higher unit price can earn its place if it consistently reduces those burdens.

The next documentation to check is concrete: Argon’s public model identifier, input-context limit, reasoning-token billing, cache terms, rate limits, supported API features, service-level availability and the introductory price end date. Those details will determine how the promising launch results translate into an application budget and a reliable workflow.

Sources and reporting notes

This article uses official provider announcements, model documentation and benchmark operators’ own pages. Numerical differences and hypothetical token bills are Kingy.ai calculations. We did not conduct hands-on Argon tests, pay for model calls or independently verify Google’s private internal deployments. Public benchmark figures can change as operators update their tables.