// Resources
The question a buyer asks after the system is already running. The four numbers that settle it, how to capture a baseline you can defend, what counts as a running cost, and how to tell time saved from time moved somewhere else.
Measure four things and the answer stops being a matter of opinion: what the work cost before, what the automated version costs to run, what the exceptions cost, and what you did with the capacity you got back. The first one is the problem, because almost nobody captures it before the build starts and nobody can reconstruct it afterwards. Self-reported time savings are the least reliable number in the room. A randomised trial run by METR in July 2025 found that experienced developers took 19% longer with AI tools while believing the tools had made them 20% faster. Count hours from system timestamps and sampled observation instead, hold the running costs against them, and re-check quarterly.
You measure whether an AI automation is saving money with four numbers, and you need all four or the answer is a guess.
The baseline. What one unit of the work cost before anything was automated, in minutes of touch time and in money, at a fully loaded rate rather than a salary divided by the hours in a nominal year.
The run cost. What the automated version costs per month to operate: model and API usage, platform fees, the human review time it still needs, and the maintenance that keeps it alive.
The exception cost. What the cases the system cannot finish cost to handle, including the ones a person has to unpick after the system got partway through. Exceptions are where most of a projected saving ends up.
The capacity outcome. What the business did with the hours. More jobs booked, more invoices processed per person, faster quotes out, a role not backfilled. Hours that come back and get absorbed into the day are real for the person and invisible in the accounts.
Subtract the second and third from the first across your monthly volume, then check the fourth to see whether the saving turned into anything. Most measurement arguments inside a company are actually arguments about the baseline, which is why we capture it during the strategy work rather than after the build, and why the number goes in writing before anybody commits to a system.
Most companies do not do any of this. The Thomson Reuters Institute's 2026 AI in Professional Services Report, built on more than 1,500 professionals across legal, tax, accounting, risk and government work, found organisation-wide AI adoption at 40%, close to double the 22% of a year earlier. Only 18% said their organisation tracks return on investment on AI at all, mostly through internal metrics, and another 40% did not know whether anybody measures it.
Measure six things about the current process before anything gets built, captured over two normal weeks rather than reconstructed from memory.
Volume. How many units a month: invoices, intake calls, quotes, confirmations, reports. Count from a system of record, not from an estimate.
Touch time per unit. How long a person actually spends, measured by watching a sample rather than by asking. People underestimate routine work they do on autopilot and overestimate work they dislike.
Cycle time per unit. How long from arrival to finished, including the hours it sits in a queue. Cycle time is usually where the money is, because a quote that goes out in an hour wins work a quote that goes out in three days does not.
Error and rework rate. What fraction comes back, and what fixing one costs. A process with a real rework rate has a second process hidden inside it, staffed by the same people.
Fully loaded hourly cost of whoever does it. Salary plus payroll taxes, benefits, software seats, supervision and the share of overhead that person carries. The Bureau of Labor Statistics' Employer Costs for Employee Compensation survey put total employer cost for private industry workers at $46.15 an hour in December 2025, of which wages and salaries were 70.1% and benefits the other 29.9%. Wages alone understate the hour before you add seats, supervision and overhead, so a process staffed by a coordinator on $28 an hour does not cost $28 an hour.
Who else is affected. The person who chases the missing information, the manager who approves, the client who waits.
An external benchmark is useful for sanity, never as a substitute. Ardent Partners' Accounts Payable Metrics That Matter in 2025 puts the average cost of processing one invoice at $9.40 against $2.78 for best in class, and average processing time at 9.2 days against 3.1 days. Ardent's separate State of ePayables 2025 research counts labor, overhead and technology together and lands higher, at $10.89 an invoice fully loaded. Knowing which of those two somebody is quoting at you matters, because they measure different things. If your own number lands nowhere near either, recheck your measurement before you use it to justify anything.
The real cost of running an AI automation is everything that would stop if you switched the system off, which is more lines than a vendor quote shows.
Model and API usage, metered. Usage scales with volume and with how much context each call carries, so a system that looks cheap in a pilot on forty documents a month behaves differently at four thousand.
Platform and infrastructure. Where the system runs, what it stores, what it connects through.
Human review. Most systems that touch money or patients or contracts keep a person in the loop by design. Review time is a cost of the automated process, not a leftover from the old one, and it belongs in the sum.
Exception handling. The share of cases the system routes to a person, times what those take. Exceptions cost more per unit than the original manual work did, because the person is picking up a half-finished job rather than starting clean.
Rework caused by the system. Output that looks finished and is not. BetterUp Labs and Stanford's Social Media Lab surveyed 1,150 US full-time workers for Harvard Business Review in September 2025 and found 40% had received AI generated work that looked plausible and failed to advance the task, each instance taking an average of one hour and 56 minutes to sort out, which the researchers costed at about $186 per worker per month.
Maintenance and operation. Monitoring, vendor migrations when an API retires, fixes when a document layout changes. Systems BLACKSIG builds live on our infrastructure and we run them from there, so this line sits with us rather than appearing as a surprise on your side. What the work involves either way is in who maintains an AI automation.
Compare the end to end cycle time, not the step you automated, and count the work that appeared downstream.
Time moves in three predictable ways. It moves to review, when a person now checks output instead of producing it. It moves to exceptions, when the system takes the straightforward majority of cases and leaves the hard remainder that used to be spread across the day. And it moves to repair, when somebody fixes output that arrived looking finished.
None of those make automation pointless. Review of 300 invoices is cheaper than keying 300 invoices. What matters is that all three show up in the measurement, because a step level saving with a worse end to end cycle time is not a saving.
Self-reported numbers will not catch this. METR's randomised controlled trial of experienced open source developers, published 10 July 2025, allowed or disallowed AI tools at random across 246 real tasks in repositories the 16 developers already maintained. They took 19% longer with the tools, and afterwards still estimated that the tools had sped them up by about 20%. The study measures the tools of early 2025, which is how METR scoped it, and the perception gap is the part that travels: the people doing the work are not a reliable instrument for timing it.
Three ways to capture hours, in ascending order of how much you can trust them:
| How you capture it | What it is good for | What it gets wrong |
|---|---|---|
| Ask the team | Finding which steps feel worst, before you measure anything | Estimates drift by large margins in both directions and cannot be audited |
| Timed observation on a sample | A defensible baseline when no system records the work | Costs a day or two of somebody's attention, and the sample has to be normal weeks |
| System timestamps | Ongoing measurement once the work runs through software | Only measures what the software sees, so offline steps stay invisible |
Use the second to set the baseline, then the third to track it. Keep the first for deciding where to look.
Capacity numbers, mostly, and that is the honest version of the answer.
Three kinds of result reach the accounts. Output per person rises: invoices processed per AP clerk, jobs dispatched per coordinator, matters opened per paralegal, calls answered without a second hire. Revenue rises because speed won work that slowness was losing, which is measurable by comparing win rates on quotes sent inside an hour against quotes sent the next day. And a cost line falls because something external stopped: an answering service, a contractor, an overtime bill, a per-document fee.
Two kinds do not reach the accounts and get presented as though they do. Hours given back to salaried staff who then do other work are real and valuable and do not change the payroll line; the gain shows up later as throughput, if you track throughput. And avoided hiring is only a saving against a role you were genuinely about to fill, with a requisition to point at.
Measure by segment rather than by average, because averages hide the shape of the gain. Erik Brynjolfsson, Danielle Li and Lindsey Raymond studied the staggered rollout of a generative AI assistant across 5,172 customer support agents in Generative AI at Work, published in the Quarterly Journal of Economics in May 2025. Productivity measured as issues resolved per hour rose 15% on average, with a 30% rise among the least experienced and least skilled agents and little effect on the most experienced. A single company-wide average would have hidden both halves of that.
It is also why we report what each system produced rather than a headline percentage. A percentage cannot be checked by the person reading it. The figures behind it can.
A measurement plan is one row per metric, a named owner and a cadence, agreed before the build rather than after.
| Metric | Where the number comes from | Cadence | What it tells you |
|---|---|---|---|
| Units processed | System of record | Monthly | Whether volume moved, so savings are not just a quiet drop in work |
| Touch time per unit | Timed sample at baseline, then system timestamps | Baseline, then quarterly | The operating saving per unit |
| End to end cycle time | Timestamps from arrival to finished | Monthly | Whether time moved rather than disappeared |
| Exception rate and cost | Queue counts plus a timed sample of exception handling | Monthly | The real ceiling on the saving |
| Rework rate | Items returned or corrected after being marked done | Monthly | Whether output quality is holding |
| Run cost | Vendor invoices plus metered usage | Monthly | The denominator |
| Capacity outcome | Output per person, win rate, headcount plan | Quarterly | Whether the saving turned into anything |
Set target ranges rather than single numbers, and write down in advance what result would make you stop. A plan that cannot produce a disappointing answer is not measurement.
The reason to agree this before anything gets built is that the baseline is only available before. Measurement is part of the roadmap work on every engagement we take, which is covered in what an AI roadmap is, and the same discipline is what separates the projects that survive from the ones described in why AI automation projects fail. For how big the saving tends to be once it is measured, rather than how to measure it, see how much time and money AI automation saves.
Re-measure monthly on the operating metrics and quarterly on the capacity question, and treat bad numbers as information rather than an argument.
Three outcomes are worth planning for. The system works and the saving is smaller than projected, usually because the exception rate is higher than the pilot suggested. Fix the exception path first, since that is where the money went. The system works and nothing changed in the accounts, which means the hours came back and got absorbed; decide deliberately what the capacity is for, or accept that the gain is comfort rather than output. And the system is not working, in which case stop, measure why, and be willing to turn it off.
That third outcome is common enough to plan for. MIT's Project NANDA report The GenAI Divide: State of AI in Business 2025, built from 52 structured interviews, 153 surveyed leaders and a review of more than 300 publicly disclosed AI initiatives, found about 95% of generative AI pilots produced no measurable profit and loss impact. The measure there is P&L impact rather than whether anybody liked the tool, so read it as a statement about how pilots are chosen and measured rather than about what the technology can do.
We decide what is worth building from the numbers the strategy work produces, our engineering team builds what the plan calls for, and we report what each system produced so the measurement does not depend on taking our word for it. The system running at Universidad Maimonides gave a team about 15 hours a week back and has been running ever since, on our infrastructure.
Measuring costs a few days of attention spread over two weeks, mostly somebody timing a sample of real work and pulling counts out of your systems. The expensive version is skipping it, then arguing for a year about whether a system that already exists is worth keeping. BLACKSIG does not publish a rate card, and the measurement work is part of the strategy engagement rather than a separate line. The method for working out what the manual version is costing you is in how much AI automation costs.
Watch it. Pick two normal weeks, time a sample of real cases end to end, count the units from whatever system touches them, and write down the rework you see. Two weeks of observation beats a year of estimates, and it gives you a baseline you can defend when somebody questions the result later. Where the work leaves no trace at all, such as phone calls handled and never logged, start by logging it for a fortnight before measuring anything.
Yes, in three ways. Usage is metered, so the cost line moves with volume instead of sitting flat. Output quality varies by case rather than being right or wrong, so you need a rework rate and not just an uptime figure. And the gain concentrates in your least experienced people, as the Quarterly Journal of Economics study of 5,172 support agents found, so a single average hides what happened. Everything else works the way any operations measurement works.
Yes, and that is most processes. Email and calendar systems carry timestamps, spreadsheets carry edit history, phone systems carry call records, and the gaps between them are exactly where the cycle time hides. We map the process as it actually runs, including the parts that live in somebody's inbox, which is the first thing the strategy work does.
Then the hours came back and nobody decided what to do with them, which is a management decision rather than a technical failure. Pick one: take on more volume with the same team, move people onto work you have been postponing, or stop backfilling a role. The point of capacity is that it lets you say yes to more, and capacity nobody spends looks identical to no capacity at all.
We map how your business actually runs, decide where AI is worth using, then our engineering team builds what the plan calls for and we run it from there. We own the outcome, not the deliverable.