How do you measure the starting point?
A baseline is only as good as the runs it is based on. Pick a representative selection of ordinary work, not just the easiest or the worst cases, and record every stage: preparation, processing, review, corrections and the handover to the next person or system. Teams often time only the step they intend to automate and forget the checking and rework around it, which is where much of the effort sits. That is why business process automation starts with the whole process, not one step.
Make the unit of work explicit. “One validated enquiry” or “one approved document” is something you can count before and after the pilot; “time spent on enquiries” is not. With a fixed unit you compare like with like, even when the volume changes from week to week. For a small team, small business automation often starts with exactly this kind of measurement.
Record volume and exceptions separately too. If the pilot period happens to bring simpler cases, or a run of unusual ones, the average time per unit will shift for reasons that have nothing to do with automation. Keeping the case mix visible lets you spot this. And avoid the classic shortcut of timing the fastest example with a stopwatch: one quick run by your most experienced person is not a baseline for the whole team.
What belongs in the operating scenario?
Once the baseline is in place, estimate the recurring change in time from the observed volume and the effort per unit before and after. Base this on measured runs from the pilot, not on a vendor’s demonstration. Then list the recurring costs that come with the new workflow: usage charges from the AI provider, software subscriptions, maintenance and, often forgotten, the time people spend reviewing the output.
Keep setup and migration work visible as a separate line. Building the integration, preparing data and training staff are real costs, but they behave differently from monthly running costs, and mixing the two makes the result hard to interpret. Anyone reading the report should be able to see both: what it took to start and what it costs to keep going.
Be careful with the word “savings”. When people spend less time on a task, you release capacity, and you can put a value on it using a cost assumption the business has signed off (our automation savings calculator does this with your own figures). But released hours do not turn into cash on their own. They become a saving only when someone makes an operational decision that reduces expense, or when the time goes to work that brings in revenue. The report should say which of these is planned and label everything else as capacity.
Why does quality decide whether the pilot passes?
Speed is easy to measure; quality takes more discipline. Track the types of error, how serious each one is, how long corrections take and what share of cases has to be escalated to a person. Review the same mix of tasks before and after the pilot, otherwise any difference in quality may simply reflect different work.
Pay particular attention to where errors end up. A workflow that is faster at the first step but passes mistakes downstream, to accounting, to the warehouse or to the customer, can create costs that are larger than the time it saves and much harder to see. This is why the full cycle up to an approved, correct result matters more than the speed of any single step. Operations analytics shows where in that cycle work gets stuck.
Write down the acceptance criteria before the pilot starts, so that the decision is not shaped by whichever number looks best afterwards. After a period of real operation, revisit your assumptions about volume, review effort and costs. What you are aiming for is a decision backed by evidence (continue, change or stop), not a single attractive payback figure. If you want an outside view before committing, AI consulting can help you set these criteria.
Checklist: measuring a pilot
- Use a stable unit of completed work, such as one validated enquiry or one approved document.
- Include review, correction and exceptions in the time you measure, not only the automated step.
- Show setup costs, recurring costs and your assumptions separately in the report.
- Review the change in time alongside the severity of errors and the quality of the handover.
Example: a fast draft, but the same review
With the pilot in place, preparing a draft takes seconds instead of several minutes. On paper this looks like a dramatic gain. But the reviewer still checks every record, and the review takes as long as it did before. The pilot report therefore measures the full cycle up to an approved record rather than the draft step alone, and the real change turns out to be much smaller than the headline figure suggests.
The browser calculator on this site can model released capacity from your own inputs. It does not include every project cost and does not predict cash savings, so treat its result as one input to the decision, not as the decision itself.


