Skip to content

Playbooks · Verification

The hours AI saves are going on checking it. Spend them where a wrong answer costs most

The hours AI saves are going on checking its output, and the professionals who keep the gain tier their checks by what breaks if an answer is wrong.

by ·4 min read·

New here? Start with the free AI Survival Kit →

The hours AI saves are going on checking it. Spend them where a wrong answer costs most

0:00 / 6:42

This voice is generated by AI.

The hours AI saves are going on checking it. Spend them where a wrong answer costs most

The productivity case for AI counts the hours saved and stops there. This year's surveys count the other column. Glean's Work AI Index, reported by CIO Dive in June, found workers save about 11 hours a week with AI, then spend nearly six and a half of them on maintenance: feeding agents context, checking output, flagging mistakes, cleaning up answers. Foxit, cited by Accounting Today, found US respondents finished the week 10 minutes down once validation time was counted. And a September survey of 500 US workers by Kolmogorov Law found only 35% always verify an AI answer before acting on it, while 51% believe they would personally carry the legal responsibility if a wrong one caused financial harm.

The check is where the job went

The consensus treats verification as overhead to be driven down. Sage, the accounting software company, called it "the verification tax" after finding that 48% of finance professionals spend 15 or more hours a week on verification activities, and 19% spend more than 30.

Read the same numbers the other way. MIT Sloan research cited by Accounting Today states that AI makes it cheap to produce work "but not to judge whether that work is any good." Production is the routine layer. Judgement is what is left, and it is what gets paid. In the same Sage survey, 71% of finance leaders said they would veto a 99%-accurate tool that could not produce a human-readable reasoning trace for every decision. The people who own the outcome are refusing output they cannot check.

AI does the routine work. You do the thinking. Verification is the thinking, done against a deadline.

Checking harder does not fix it

The awkward finding in the Kolmogorov Law survey concerns the careful group. Of the 173 respondents who said they always verify, 40% had still accepted an answer they suspected was wrong. They reported work problems at 32%, against 30% for the whole sample. Ticking "always" bought no protection. The sample is small and unweighted, and the pattern matches the lab.

Wharton researchers ran three experiments with 1,372 participants. When the AI gave a confident but incorrect answer, 73% accepted it, and using AI raised their confidence in the wrong answers. The researchers call this cognitive surrender. The answer arrives fast and reads smoothly, and the checking stops without anyone deciding to stop it.

Two conclusions follow. A heavy check on everything is where the 15-hour weeks come from. A light check on everything is theatre. The risk is uneven: of the 113 respondents who use AI for legal, financial or compliance questions, 43% said a wrong answer had already cost them something, and 64% of that group still do not always verify.

Unchecked output also travels. BetterUp and the Stanford Social Media Lab found 40% of US full-time employees had received AI "workslop" in the previous month, polished content that does not advance the task, and 18% of it came from direct reports to their managers. Skipping the check moves the cost to whoever reads your work next. They remember who sent it.

The operator response is to spend checking time where the cost of error sits, and nowhere else.

Your Next Move

  1. Keep a two-week ledger. Each time AI produces work output for you, note the minutes it saved and the minutes you spent checking, correcting or redoing. Glean's averages say nothing about your role; your own ledger shows which tasks pay and which quietly run at a loss. Redesign or drop the losers. Glean found more than a third of AI sessions fail completely, so expect some.

  2. Sort every AI use into three tiers by what breaks if the answer is wrong. Tier one covers internal drafts and thinking aids: read for sense and move on. Tier two covers anything that leaves your hands for a colleague: check every figure and factual claim against a source you opened yourself. Tier three covers legal, financial, compliance and client-facing work: verify against primary sources and record what you checked. Do not hand the check to a second model. The MIT paper warns that two systems sharing the same assumptions reinforce the same errors.

  3. Write your verification standard down and send it to your manager. Only 22% of Kolmogorov Law respondents said their employer has a written policy requiring AI output to be verified; 59% said there is none. One page is enough: the tiers, the checks, what gets logged. It puts your judgement on record before the next review, and when a wrong answer does get through, the record shows the process you ran.

I build systems like the one publishing this site. → Work with me

About the author

Jo

Jo runs The War Room: one signal a day on how AI is changing work, and what to do about it.

Sources

Get the Briefing

Want more intelligence like this?

The weekly briefing, free — and the AI Survival Kit with it.

Takes 30 seconds. No spam.

More in Playbooks

See all