Four numbers convince a partner group: hours per instance, cycle time, rework rate, and realization. Capture all four before the pilot starts, measure them again at the end, and you have an argument. Skip the baseline and you have an anecdote, which is what most AI pilots produce and why most second-round funding requests fail.
I have never seen a firm struggle with the arithmetic here. The hard part is not the math. It is that the measurement has to happen before anyone is excited, when nobody wants to spend a week counting hours on a workflow they already believe is inefficient.
Why the usual numbers fail in a partnership
The figure most AI pilots report is time saved per task, sourced from the people who ran the pilot. It never survives a partner meeting, and it should not. It is self-reported by advocates, it counts the good instances, and it cannot be reconciled against anything in your accounting system.
Partners are not being obstructive when they reject it. They are applying the same skepticism they would to any capital request with no verifiable basis. The fix is to measure things that already exist in your systems, where the pilot team cannot influence the number.
The four numbers
1. Hours per instance
Direct labor hours to complete one unit of the workflow, one submittal review, one report section, one proposal. Pull it from timesheets rather than asking people, and average across at least ten instances so a single unusual project does not set your baseline.
The trap is scope creep in the definition. "Hours per submittal review" has to mean the same activity before and after, or your improvement is measurement drift. Write the definition down in one sentence before you collect anything, and hold to it.
Expect this number to get worse for the first two to three weeks of the pilot. That is learning cost, it is real, and reporting it honestly buys you credibility for the rest of the numbers.
2. Cycle time
Calendar days from request to delivered. Almost always the number clients actually experience, and almost always far worse than the labor hours suggest, because most of a submittal's life is spent in someone's queue rather than under someone's attention.
This is frequently where AI produces its largest and least expected gain. A first-pass review that used to wait three days for an available engineer now happens the same afternoon, and the total cycle drops by more than the hours saved would predict. It is also the number your clients notice, which makes it the most useful one in a business development conversation.
3. Rework rate
How often the work comes back from review needing substantive correction. The essential counterweight to the first two numbers, because a workflow that gets faster and sloppier is not an improvement, and without this number your pilot cannot distinguish the two.
Count from your existing QA records if you have them. If you do not, start counting now, a simple tally of "returned for substantive correction" against total instances is enough, and the absence of any such record is itself a finding on the readiness assessment.
Watch this one hardest. If rework rises, you have a training gap rather than a tool problem, and the answer is teaching people the specific failure modes rather than abandoning the pilot.
4. Realization
The number that makes partners lean in, because it is already on your financial reports: billed value against the labor cost of delivering it, on the projects touched by the pilot.
Realization is where efficiency becomes money, and it is where most AI business cases quietly break. If you save twelve hours on a lump-sum project, realization improves. If you save twelve hours on a time-and-materials project and simply bill less, realization is flat and the firm has gained nothing. You have handed the savings to your client. Same tool, same hours saved, opposite financial result.
Which is why the question of what happens to recovered hours has to be answered before the pilot, not after. It is a commercial decision, not a technology one.
What to do with recovered hours
Three options, and you should choose deliberately rather than letting it happen:
- Take on more work. The strongest option if you have demand you have been declining. Converts saved hours into revenue at full margin, and it is the easiest case to make to a partner group.
- Reduce overtime and burnout. Harder to put on a spreadsheet, real on your retention numbers. Worth quantifying against what a departure actually costs you in recruiting and lost project continuity.
- Reinvest in quality or pursuit work. Better QA, more thorough proposals, more time on the technical approach. Shows up in win rate rather than utilization, on a longer lag.
What you must not do is leave it undeclared. Staff will assume the answer is "we bill less and eventually need fewer of you," and adoption will quietly stall, for entirely rational reasons.
What not to measure
Tokens, prompts, queries, logins. Vendor metrics. They describe activity, not value, and a firm reporting query volume to its partners has substituted motion for results.
Satisfaction surveys. Everyone reports satisfaction in month one. Nothing predictive.
Projected annual savings extrapolated from a good week. The single fastest way to lose credibility. Take the measured result across the full pilot, including the bad first fortnight, and present it with the range.
Bring a modest number you can stand behind. An impressive one falls apart on the first question, and then nobody trusts the next one either.
The one-page report
When the pilot ends, the partner conversation should need one page: the workflow in a sentence, the four numbers before and after, the licenses and internal hours it consumed, what happens to the recovered capacity, and a recommendation to scale, adjust, or stop.
Include the cost side honestly, including the internal time, the pilot lead's hours, the training, the IT work. Those are usually larger than the license fee and their omission is what makes a business case look manufactured. The full picture is in what AI actually costs an engineering firm.
And be willing to recommend stopping. A pilot that concludes "this did not pay off on this workflow" is a successful pilot. It cost you twelve weeks instead of a firm-wide rollout, and it makes the next recommendation believable. Firms that never report a negative result train their partners to discount every result. If yours did fail, the diagnosis matters: most failures have four causes, none of them technical.
If you want the baseline set up properly before you start, or a pilot result put into a form your partners will accept, that is what our engagements are built around, every one ends with a number. Start a conversation.
← Back to insights