Signals · Capability Displacement
Google built Gemini 4 Argon for legal, finance and coding work, and keeps people reviewing its critical rewrites
Google built Gemini 4 Argon for legal, finance, tax and coding work. Its large rewrites of critical code still get human audits before they ship.
by Jo·3 min read··Updated
New here? Start with the free AI Survival Kit →
Google built Gemini 4 Argon for legal, finance and coding work, and keeps people reviewing its critical rewrites
0:00 / 6:30
This voice is generated by AI.
Part of the guide Will AI take your job? How to find out for your role, in an afternoon
Google announced Gemini 4 Argon, its new frontier model, on its official blog on 30 September. The post was written by Koray Kavukcuoglu, SVP of Google DeepMind and Chief AI Architect. Argon is going first to a set of trusted cyber defenders through Google's Fairwind Program. Google names three targets: real-world software engineering, enterprise knowledge work "like legal and finance", and cybersecurity defence. Developers, enterprises and consumers follow "as soon as possible", once Google has iterated on guardrails with its early testers. The introductory price is $2 per million input tokens and $10 per million output tokens.
Google's chosen benchmark shows the target
Google says Argon leads the Vals Index. That index measures economic impact across finance, coding, legal and tax work, and weights every sector by its contribution to US GDP. So the scoreboard here is weighted by the economy's paid work.
The other numbers follow the same pattern. On DeepSWE v1.1, which tests long-horizon software engineering, Argon scores 77.9%. On AutomationBench, Zapier's test of end-to-end execution across core business functions, it ranks first with 51.3%. Google also raised the output token limit to 1M, up from 64K, so the model can work through a long, multi-step task in a single run.
The ranking is the headline. The number underneath it is more useful: 51.3% on end-to-end execution across core business functions is the lowest benchmark score Google reports for Argon. Google does not say how that score maps onto a real business process. Where a model finishes about half of an end-to-end test, someone has to check the output and own the result. That check is the part the benchmark cannot score.
How Google uses it itself
The clearest evidence is how Google uses Argon itself. Google says thousands of its staff use the model for specialised coding, deeper research and writing. A team of Argon agents studied fleet-wide profiling data and, on their own, applied memory optimisations across Google's data centres. Once rolled out, these freed more than 300 TiB of memory. On the libgav1 video decoder, Argon agents replaced 32K lines of SIMD code. The result was a memory-safe decoder that runs 2.7x faster than the earlier Rust port.
Now look at the large codebase migrations. Argon agents are moving C/C++ code to Rust, up to 800K+ lines for the Fuchsia Zircon kernel. Google says these rewrites go through "rigorous automated and manual auditing, emulation testing, and review" before they reach production.
Google gives the reason as the criticality of those systems. The memory work ran autonomously; the kernel and library rewrites get people auditing, testing and reviewing them. Google ties the human review to how critical the system is. AI does the routine work. You do the thinking. On critical systems, the review seat is where accountability sits.
The second signal is about access. Google is giving Argon to trusted defenders and its own internal teams without cyber guardrails. It is also taking part in the US government's voluntary pre-release access process while it widens access step by step. Separately, Wiz is using Argon through its Scan for Good initiative and found a critical vulnerability exposing sensitive personal information in healthcare software used by hospitals worldwide. Google says earlier frontier models had missed it.
Your Next Move
-
Sort your week against the Vals sectors. If your work involves finance, legal, tax or code, list last week's long, document-heavy, multi-step tasks. Argon was built for those. For a step-by-step version, see the method for sorting your role's tasks and scoring your exposure.
-
Take the review seat before someone else does. Offer to own sign-off on AI-produced work in your team: the audit, the test plan, the go or no-go decision. Google keeps that layer staffed for its critical rewrites. You should hold it in yours.
-
Keep a failure log. When Argon becomes available to you, run one real task from your own work through it. Write down exactly where it fails and why you caught it. That record is your judgement in writing, and it is the evidence you bring to your next review.
I build systems like the one publishing this site. → Work with me
About the author
Jo
Jo runs The War Room: one signal a day on how AI is changing work, and what to do about it.
Sources
Get the Briefing
Want more intelligence like this?
The weekly briefing, free — and the AI Survival Kit with it.
More in Signals
See allSignals · Capability Displacement
Utah's AI prescribing pilot shows how a profession hands over a task in three phases
3 min read
Signals · Governance
OpenAI's safety lead quit over culture, and his fix is the human oversight every team running AI agents now needs
3 min read
Signals · AI Adoption
A two-year Khanmigo trial shows that access to AI is not the same as using it
3 min read