ELU NOVA
← All field notes
The pilot-to-production gap

Your POC worked because it never met the exception path

4 min read

Two different workloads wearing the same name

A proof of concept is usually run on a curated sample: documents that scanned properly, cases that follow the standard route, records that exist in the system they are supposed to exist in. This is not dishonesty. It is how you isolate the question of whether the technology can do the task at all.

Production is a different distribution. It contains the scan taken at an angle in a warehouse, the supplier who changed their template without telling anyone, the record that exists twice, the case that was already half-processed manually before the automation saw it, and the transaction that is legitimate but does not resemble anything in the sample.

The technology question was answered by the POC. The economics question was not, because the economics are determined almost entirely by the part of the distribution the POC excluded.

Gartner, in its 25 June 2025 assessment, named inadequate risk controls among the three causes of the agentic project cancellations it predicted through 2027 — alongside escalating costs and unclear business value. MIT's GenAI Divide, published the same month, attributed widespread stalling to systems that could not retain feedback, adapt to context or improve over time. Both descriptions are about what happens after the clean path ends.

The arithmetic of the exception queue

Assume a workflow at 10,000 documents a month. The POC suggested a 90% straight-through rate.

That leaves 1,000 exceptions a month. At eight minutes of handling each — a realistic figure once the operator has to open the source document, work out what went wrong and correct it — that is roughly 133 hours a month, about three-quarters of a full-time role.

The saving is still real: 9,000 documents that no longer need touching. But the business case that assumed near-total automation was wrong by most of a headcount, and the operating model that assumed the existing team would absorb exceptions "as part of their normal work" has quietly created an unowned queue.

Now the part that determines whether the deployment survives its second year. If exceptions arrive at 1,000 a month and are cleared at 800, the queue grows by 200 a month indefinitely. Nobody notices for a quarter. Then someone notices, and the automation is blamed for a backlog it did not create.

Exception arrival rate versus exception clearance rate is the single most important operational number in any automation deployment, and it is almost never in the business case.

Sizing it before you commit

Sample the real distribution, not the clean one. Pull a genuinely random month of inputs, including the ones people complain about. If the POC sample was assembled by someone helpful, it is not random.

Measure the current manual exception rate. How often does a human already have to do something unusual with this workflow? The automation will not have a lower exception rate than the process does; it will surface the exceptions that were previously absorbed invisibly by experienced staff.

Cost handling time honestly. Time an operator working real exceptions in the actual interface. Handling time estimated by a manager is consistently optimistic, usually by a factor of two.

Ask the vendor for their production exception rate on comparable documents. Not accuracy. Exception rate, at volume, on inputs like yours. A vendor without that number does not have production deployments in your situation.

Designing so the queue drains

Name the owner. One person accountable for exception queue depth, reported weekly. Unowned queues grow.

Set a depth alarm, not a rate target. A rising rate can be benign; rising depth never is. Alarm on days-of-backlog.

Make corrections feed back. If an operator's correction changes nothing about future behaviour, your exception rate is fixed forever. MIT's learning-gap finding is precisely this. The feedback path is engineering work and it is worth budgeting.

Build the review interface for speed. The source document, the extracted value, the confidence, the rule that flagged it — on one screen, with a keyboard path. A reviewer who has to open three systems to resolve one exception will approve without checking, which is worse than no review at all.

Triage by consequence, not by confidence alone. A low-confidence payment term and a low-confidence bank detail should not sit in the same queue at the same priority.

The conversation to have before signing

Ask for the business case to be restated with a realistic exception rate, realistic handling time, and a named owner for the queue. If the saving still justifies the build, you have a case you can defend when someone audits it in eighteen months.

If it does not, you have learned that for the cost of a meeting rather than the cost of a programme. That is not a failed evaluation. It is the evaluation working.

Find out where AI actually pays off in your business.

A focused discovery audit with a senior principal. You leave with a prioritized, costed roadmap — whether or not you work with us.

contact@elunovatech.com · Response within one business day