Failed AI pilots in engineering firms fail for four reasons, and none of them is the model. No baseline, so the result could not be proven. The wrong workflow, so the result did not matter. No owner, so nothing changed. Or an unanswered question about what happens to the saved hours, so people declined to cooperate. Diagnose which one killed yours and the restart is straightforward. Skip the diagnosis and the second attempt fails the same way with less patience available.
I care about getting this diagnosis right because a firm gets a limited number of attempts. After two visible failures, the internal position hardens into "we tried that," and the people who would have championed the third attempt have stopped volunteering.
Cause one: no baseline, so there was no result
The most common by a wide margin. The pilot ran, people found it useful, and at the end nobody could say what it was worth. The report reads "the team reported significant time savings," which a partner group correctly treats as an opinion.
What makes this fatal is that it is unrecoverable after the fact. You cannot reconstruct last quarter's hours per submittal once the workflow has changed. The measurement window closed.
The symptom: your pilot summary contains adjectives where numbers should be.
The restart: spend two weeks measuring before you resume anything. Hours per instance from timesheets, cycle time in calendar days, rework rate from QA records. It feels like a delay. It is also the whole difference between a result and an anecdote, which is what the four numbers that convince a partner group covers.
Cause two: the wrong workflow
The pilot targeted something interesting rather than something measurable. Usually design, occasionally an infrequent task, sometimes a workflow so entangled with judgment that no honest before-and-after exists.
Design pilots fail predictably: your senior engineers become opponents because they are accountable for the output, the licensure conversation consumes the effort, and the result is unmeasurable because design quality is a matter of professional judgment. That is not a technology failure. It was decided when the workflow was chosen.
The symptom: the pilot generated more debate about whether the output was acceptable than data about whether it was faster.
The restart: pick a workflow that is weekly, countable, and distant from a seal. Submittals, RFI logs, report boilerplate, proposal assembly, minutes. The full test is in why submittals beat design as a first pilot.
Cause three: no owner whose week improved
The pilot was assigned rather than wanted. Someone with capacity was asked to evaluate a tool, did so conscientiously, produced a reasonable summary, and returned to their actual job. Nothing about their week was worse before the pilot or better after it, so there was no reason for the change to persist.
Committee-owned pilots are a variant of the same failure. A committee produces a recommendation, and a recommendation is not a changed workflow.
The symptom: usage went to near-zero within weeks of the pilot ending, and nobody complained.
The restart: find the person who is in pain. The project manager losing every Thursday to submittal logs, the coordinator reconciling document control by hand. Give it to them and let them define what good looks like. A pilot owned by someone in pain gets defended in meetings for a year; a pilot owned by someone with capacity gets filed.
Cause four: the unanswered question
The quietest cause and the hardest to see, because it never appears in the pilot report. Staff worked out that saved hours meant reduced utilization, or eventually fewer people, and declined to cooperate, not by objecting, which would be visible, but by using the tool minimally and reporting that it did not help much.
They were behaving rationally. Nobody volunteers to shrink their own billable hours in a firm that measures them on billable hours. If leadership never says what happens to recovered capacity, people assume the answer that is worst for them.
The symptom: polite, low-energy participation. Feedback that is neither enthusiastic nor critical. Adoption that looks like compliance.
The restart: answer the question before the next pilot begins, out loud, specifically. "We take on the two pursuits a month we have been declining." "You get your Thursdays back." "Early-career staff spend that time on engineering instead of formatting." Then hold to it visibly on the first pilot, because the second one will be judged by whether you did.
Two causes that get blamed and usually are not it
"We picked the wrong tool." Occasionally true, mostly a comfortable explanation because it implies the fix is a purchase. Test it honestly: was there a baseline? Did the workflow qualify? Did anyone own it? If any answer is no, a different tool would have failed identically.
"Our people resisted." Sometimes a training gap misread as resistance. Engineers who have not been taught the failure modes get burned by a fabricated citation, conclude the tool is unreliable, and stop, which is a reasonable response to their actual experience. That is not resistance; it is an untrained reviewer behaving correctly. It is fixed by teaching the five specific failure patterns rather than by change management.
How to restart without burning the remaining patience
Say plainly that the first one did not work, and why. Firms that quietly rebrand a failure lose credibility with exactly the people whose support they need. A specific diagnosis, "we had no baseline, so we could not prove anything", is far more persuasive than a fresh round of optimism.
Make the second attempt smaller. One workflow, one team, twelve weeks, one number. The instinct after a failure is to go bigger to demonstrate seriousness. Go narrower instead; a small proven result buys more than a large ambiguous one.
Commit to reporting the number even if it is bad. This is what makes the third attempt possible. A firm that has once reported "this did not pay off on this workflow, here is the measurement" has established that its results mean something.
If a pilot ends with "not on this workflow, and here is the measurement," it did its job. Ending with "everyone liked it" means you learned nothing.
If you have not run one yet
Read the four causes as a design specification. Baseline first. Weekly, countable, away from the seal. An owner in pain. The capacity question answered out loud before you start. That is most of what separates the pilots that scale from the pilots that get quietly discontinued, and none of it depends on which tool you buy.
It is also worth checking the foundation before you commit. Two of the four causes are visible in advance on the twelve-question readiness assessment, the baseline question and the senior sponsor question, and a firm scoring low on both should fix that before spending a quarter on a pilot.
If you want a post-mortem on a pilot that stalled, or a second attempt structured so it produces a defensible result, that is a useful thirty minutes. Our engagements all end with a number, which is the discipline most first pilots were missing. Start a conversation.
← Back to insights