Write the decision before the test
State what the pilot is meant to decide. For example: “Can this team reduce average human processing time for supplier purchase invoices while keeping every posted entry reviewed?” This prevents the evaluation from changing into a demonstration of whichever metric looks best.
Build a representative sample
Use at least 30 documents and retain the mix your team actually receives: electronic PDFs, scans, phone photos, multi-page invoices, repeat suppliers, new suppliers, stock-heavy lines and account-coded expenses. Do not remove failed or awkward examples after processing.
Exclude genuinely invalid documents only under a written rule, such as a file that contains no readable invoice at all. Report how many were excluded and why.
Record the manual baseline
Measure collection, data entry, code lookup, checking, posting and correction. Use the same accounting system and the same kind of operator who will use the live process. Report the median as well as the average so one unusually long invoice does not distort the result.
Measure outcomes that matter
| Metric | How to calculate it |
|---|---|
| Total human minutes per document | Collection + review + correction + delivery time divided by all attempted documents. |
| Ready without correction | Documents requiring no field or code change divided by attempted documents. |
| Review-required rate | Documents with at least one blocker or correction divided by attempted documents. |
| Successful accounting delivery | Entries accepted by SQL Account or AutoCount Cloud divided by approved documents. |
| Duplicate prevention | Repeated submissions correctly stopped or skipped during retry. |
Separate extraction errors from setup errors
A missing supplier code may mean the master data was not imported; a failed AutoCount push may mean the default purchase location is missing; a SQL Account API failure may mean all licensed database connections are occupied. Record these separately from a misread date or amount.
Publish the method beside any claim
A defensible case study states the sample size, document mix, measurement period, baseline, review policy and exclusions. It avoids phrases such as “99% accurate” unless the denominator and field-level scoring method are visible.
beres does not publish pilot outcomes on this site until those figures have been supplied and verified. The calculator is therefore labelled illustrative rather than presented as customer evidence.