Time spent checking AI output: the saving that verification eats
Time spent checking AI output cancels 40% of the saving. What finished work has to mean, and how to make verification cheap instead of endless.
Time spent checking AI output cancels roughly 40% of the time AI saves. Workday research published in January 2026 found that close to 40% of the saving goes back into rework — rewriting, fact-checking and correcting — while 85% of workers report saving one to seven hours a week. The headline saving is real; the net saving is about half of it.
How much does time spent checking AI output actually cost?
Enough to change any business case built on the headline number:
- ~40% of time saved is spent on rework: rewriting, fact-checking, correcting — Workday, January 2026
- 85% of workers save one to seven hours a week, losing part of it to validation — CFO.com on the same research
- ~40% of US desk workers received "workslop" in a single month; each incident took close to two hours to resolve, at an estimated $186 per employee per month — research popularised by Harvard Business Review
- ~40% of that burden travelled between peers, with the rest between managers and their teams
Two of those figures interact badly. If a tool saves five hours a week and returns two to rework, the remaining three are real — until one workslop incident lands on a colleague and costs them two more. At that point the team is level, and the person who feels most productive is the one who created the work.
What is workslop, and why is it worse than a bad draft?
Workslop is AI-generated work that looks finished and is not: plausible structure, confident tone, missing or wrong substance. It is worse than an obviously bad draft because the cost lands on someone else.
The person who generated the document saved twenty minutes. A colleague spent two hours working out which parts were true. Measured individually, the tool looks like a success. Measured across the team, it may be a net loss — which is why "our people say it saves them time" is a weak signal, and why the honest metric is finished work leaving the team, not drafts entering it.
Why is a draft not finished work?
Because a draft resembles the artefact and a result is the artefact. That gap is exactly the rework the studies measure, and it is invisible at a glance: a confident paragraph and a correct paragraph look identical until someone checks. That is what makes verification expensive — you cannot skip it, because you cannot tell from the surface whether this was the time you needed to.
What does "finished" have to mean?
Four properties, decided before you automate anything.
It is in the real format. The actual spreadsheet, report template or reply — not a description of what should go in one. If a person transfers content into the real artefact, you moved the work rather than removed it.
Every figure has a stated source. Not "according to the data" but which file, which system, which field, pulled when. A reviewer should be able to check one number without reconstructing the process.
Exceptions are raised, not guessed. This property decides whether verification is cheap or expensive. A system that fills a gap with something plausible forces a human to check everything, because any field might be the invented one. A system that flags three unresolved fields lets the human check three fields.
Scope was agreed before work started. If the request is restated back before anything runs, misunderstandings surface in the ten seconds before rather than the review after.
The pattern behind all four: verification cost falls when the system tells you where to look, and rises when output is uniformly confident.
Kvantia Harness is a supervision layer for an AI worker on a Windows PC your business controls, and this is the problem it is shaped around: a run returns the result together with what was read, which actions were taken, which were queued for a person, and what could not be resolved. The point is not that a record exists — it is that the reviewer knows which three fields to examine instead of re-reading everything. How it works shows a routine from request to audit trail.
How do you measure whether you are actually ahead?
Two numbers are enough.
Time from trigger to accepted result, including review. Not generation time. If a report generates in 30 seconds and takes 40 minutes to check, your number is 40 minutes and change.
Rework rate. What share of outputs came back for correction, and how long each took. If you cannot answer that after a month, the tool is not being evaluated — it is being tolerated.
Both numbers should be collected by the person doing the review, not the person who introduced the tool. That is not a trust problem; it is that the two roles see different halves of the cost, and only the reviewer sees the half the studies are measuring.
Track both against the baseline you measured before, and the argument stops being about whether AI is impressive. It becomes an ordinary operational question with an ordinary operational answer — the same discipline that separates the few successful pilots in why AI pilots fail, and the same reason to choose a first routine that is quick to verify, as in which tasks to automate with AI first.