Skip to content

Signals · Capability Displacement

Google built Gemini 4 Argon for legal, finance and coding work, and keeps people reviewing its critical rewrites

Google built Gemini 4 Argon for legal, finance, tax and coding work. Its large rewrites of critical code still get human audits before they ship.

by ·3 min read··Updated

New here? Start with the free AI Survival Kit →

Google built Gemini 4 Argon for legal, finance and coding work, and keeps people reviewing its critical rewrites

0:00 / 6:30

This voice is generated by AI.

Part of the guide Will AI take your job? How to find out for your role, in an afternoon

Google built Gemini 4 Argon for legal, finance and coding work, and keeps people reviewing its critical rewrites

Google announced Gemini 4 Argon, its new frontier model, on its official blog on 30 September. The post was written by Koray Kavukcuoglu, SVP of Google DeepMind and Chief AI Architect. Argon is going first to a set of trusted cyber defenders through Google's Fairwind Program. Google names three targets: real-world software engineering, enterprise knowledge work "like legal and finance", and cybersecurity defence. Developers, enterprises and consumers follow "as soon as possible", once Google has iterated on guardrails with its early testers. The introductory price is $2 per million input tokens and $10 per million output tokens.

Google's chosen benchmark shows the target

Google says Argon leads the Vals Index. That index measures economic impact across finance, coding, legal and tax work, and weights every sector by its contribution to US GDP. So the scoreboard here is weighted by the economy's paid work.

The other numbers follow the same pattern. On DeepSWE v1.1, which tests long-horizon software engineering, Argon scores 77.9%. On AutomationBench, Zapier's test of end-to-end execution across core business functions, it ranks first with 51.3%. Google also raised the output token limit to 1M, up from 64K, so the model can work through a long, multi-step task in a single run.

The ranking is the headline. The number underneath it is more useful: 51.3% on end-to-end execution across core business functions is the lowest benchmark score Google reports for Argon. Google does not say how that score maps onto a real business process. Where a model finishes about half of an end-to-end test, someone has to check the output and own the result. That check is the part the benchmark cannot score.

How Google uses it itself

The clearest evidence is how Google uses Argon itself. Google says thousands of its staff use the model for specialised coding, deeper research and writing. A team of Argon agents studied fleet-wide profiling data and, on their own, applied memory optimisations across Google's data centres. Once rolled out, these freed more than 300 TiB of memory. On the libgav1 video decoder, Argon agents replaced 32K lines of SIMD code. The result was a memory-safe decoder that runs 2.7x faster than the earlier Rust port.

Now look at the large codebase migrations. Argon agents are moving C/C++ code to Rust, up to 800K+ lines for the Fuchsia Zircon kernel. Google says these rewrites go through "rigorous automated and manual auditing, emulation testing, and review" before they reach production.

Google gives the reason as the criticality of those systems. The memory work ran autonomously; the kernel and library rewrites get people auditing, testing and reviewing them. Google ties the human review to how critical the system is. AI does the routine work. You do the thinking. On critical systems, the review seat is where accountability sits.

The second signal is about access. Google is giving Argon to trusted defenders and its own internal teams without cyber guardrails. It is also taking part in the US government's voluntary pre-release access process while it widens access step by step. Separately, Wiz is using Argon through its Scan for Good initiative and found a critical vulnerability exposing sensitive personal information in healthcare software used by hospitals worldwide. Google says earlier frontier models had missed it.

Your Next Move

  1. Sort your week against the Vals sectors. If your work involves finance, legal, tax or code, list last week's long, document-heavy, multi-step tasks. Argon was built for those. For a step-by-step version, see the method for sorting your role's tasks and scoring your exposure.

  2. Take the review seat before someone else does. Offer to own sign-off on AI-produced work in your team: the audit, the test plan, the go or no-go decision. Google keeps that layer staffed for its critical rewrites. You should hold it in yours.

  3. Keep a failure log. When Argon becomes available to you, run one real task from your own work through it. Write down exactly where it fails and why you caught it. That record is your judgement in writing, and it is the evidence you bring to your next review.

I build systems like the one publishing this site. → Work with me

About the author

Jo

Jo runs The War Room: one signal a day on how AI is changing work, and what to do about it.

Sources

Get the Briefing

Want more intelligence like this?

The weekly briefing, free — and the AI Survival Kit with it.

Takes 30 seconds. No spam.

More in Signals

See all