Which tasks to automate with AI first: repeatability times verifiability
Which tasks to automate with AI first, scored in ten minutes — plus where to put the human approval and what to measure before you build anything.
Deciding which tasks to automate with AI comes down to two properties multiplied: how repeatable the work is, and how quickly a person can verify the result. Get that choice wrong and the effort ends before anyone sees a number — which is a large part of why about 95% of enterprise AI pilots produce no measurable return. The routines people most want to hand over are usually the worst place to start.
Which tasks to automate with AI first?
The ones scoring high on both axes at once.
Repeatability. How often does this run, and how similar is each run? Weekly beats quarterly, because feedback arrives four times faster and you reach a verdict inside a quarter. Following a written procedure beats following whatever the person remembers, because there is a standard to compare against.
Verifiability. How quickly can a competent person tell whether the output is right? A reconciliation is highly verifiable: the numbers match or they do not. A strategy memo is not — two reviewers will disagree about whether it is any good, so you can never prove the automation worked.
Multiply, do not average. A daily task nobody can check is a machine for producing invisible errors at speed. A perfectly checkable task happening twice a year teaches you nothing before the budget conversation.
Why are the appealing candidates usually wrong?
Because work is unpleasant precisely when it needs judgement under ambiguity — the property that makes it both hard to automate and hard to verify. The messy inbox, the awkward customer situation, the thing that goes wrong monthly for a different reason: all bad first candidates.
Good first candidates are usually described as "boring but fiddly". Not hated — just repetitive, procedural and time-consuming in a way nobody has bothered to fix.
How do you score a candidate routine?
One to five on each, then multiply:
- Frequency. Daily 5, weekly 4, monthly 3, quarterly 2, ad hoc 1.
- Procedure. Written down and followed 5. In one person's head 2. Disputed 1.
- Sources. Named files and systems with stable structure 5. Whatever arrives by email 2.
- Output. A standard artefact in a standard template 5. Freeform prose 2.
- Checkability. Objectively right or wrong in minutes 5. A matter of taste 1.
- Reversibility. Nothing irreversible without approval 5. Sends or pays automatically 1.
Anything above roughly 400 out of 15,625 deserves a serious look. The threshold is deliberately imprecise — the value is in the conversation the scoring forces, which usually reveals that two people believed different things about the sources.
Good first routines: the weekly management pack assembled from the same three systems; the register that must match a source system; supplier invoice data extracted into a standard structure with mismatches flagged; recurring follow-up preparation where a person still presses send.
Poor, for now: anything with a disputed procedure; anything whose source data people already distrust; anything that sends externally or moves money without review; anything rare enough that you will forget how it was set up; anything where "good" is a matter of judgement.
Where should the human approval sit?
Immediately before the first step you cannot cheaply undo. The test is reversibility, not anxiety: rebuilding a report costs a re-run, un-sending an email to a customer costs a relationship, recovering a deleted register costs a day and a difficult conversation.
Two failure modes. Approving too early means signing off a plan rather than a result — the person confirms intent, the work happens afterwards, and nobody looked at the output. Approving everything turns the reviewer into a rubber stamp within a fortnight; if a person must confirm eleven low-stakes steps to reach the one that matters, they stop reading by step three.
One meaningful gate at the irreversible step beats eleven ceremonial ones.
Kvantia Harness is a supervision layer for an AI worker on a Windows PC your business controls, and this is the decision it makes explicit: you grant the capabilities a routine needs, choose the step that waits for a person, and get a record of what was read and done. Use cases walk through specific routines with their sources, approval point and what to measure.
What should you measure?
Agree these before anything is built:
- Time from trigger to accepted result, including review.
- How often output is returned for correction, and why.
- How many exceptions are raised, and whether they were genuine.
- Whether it completes on time without someone chasing it.
Then measure the current process against the same four for a few cycles. The baseline is the step everyone skips and the only reason you will be able to answer "did it work" with something better than an opinion. Review time is not a detail here — time spent checking AI output eats roughly 40% of the saving when the output is not built to be verified.
Why start narrow on purpose?
One routine, one owner, one clear gate, one measured baseline, proven over a quarter. Breadth is what makes results unattributable, and unattributable results are what killed most of the pilots described in why AI pilots fail.