Start where the work is repetitive, document-heavy, and nowhere near a seal. In most engineering firms that means submittals, RFI logs, proposal assembly, meeting minutes, or report boilerplate, not design. The first pilot should be boring enough that nobody argues about it and measurable enough that nobody can dismiss the result.

This advice disappoints almost everyone I give it to. Design is where the interesting problems are, and it is where leadership imagines AI creating advantage. But a first pilot has a job other than being interesting: it has to produce a defensible number and survive contact with senior technical staff. Design pilots fail both tests, for reasons that have nothing to do with the technology.

Why the exciting pilot is the wrong first pilot

Put AI anywhere near design and three things happen at once. Your most senior engineers, whose support determines whether anything spreads, become adversaries rather than allies, correctly, because they are responsible for the output. The licensure question arrives immediately and consumes the conversation. And the result becomes unmeasurable, because design quality is a matter of judgment and you cannot put a number on judgment in twelve weeks.

Meanwhile the actual hours in your firm are not in design. They are in the documentation surrounding it. Look at where an experienced project manager's week goes: chasing submittal status, writing the same three paragraphs of report background, assembling a proposal from four prior proposals, reconciling an RFI log, converting meeting notes into something distributable. None of it requires engineering judgment. All of it requires someone competent, and competent people are what you are short of.

That is the arbitrage. The unglamorous work is where the recoverable hours sit, and it is the work your staff will thank you for automating.

The four tests a first pilot has to pass

Score any candidate workflow against these before committing. A workflow that fails two is a bad first pilot no matter how appealing it looks.

  1. Repetition. Does it happen at least weekly? A quarterly task cannot generate enough instances to prove anything inside a pilot window.
  2. Measurability. Can you count hours or days on it today, or at least by next week? If the output is a matter of taste, you have no result, only opinions.
  3. Distance from the seal. Does the output pass through engineering judgment before it reaches a client? Good. Does it go out under a seal without that? Wrong pilot, for now.
  4. An owner who wants it. Is there a specific person whose week improves if this works? Pilots owned by a committee produce reports. Pilots owned by someone in pain produce change.

Test four is the one firms skip and the one that most often kills the effort. A pilot assigned to whoever has capacity will be executed dutifully and abandoned the moment the pilot ends. A pilot handed to the project manager who loses every Thursday to submittal logs will be defended by that person in every meeting for a year.

The four workflows that usually win

Submittal review, first pass. A tool compares a product data sheet against the relevant specification section and flags apparent discrepancies for an engineer to adjudicate. It does not approve anything. It sorts a stack into "probably compliant" and "look at this," which is where most of the time goes. High repetition, easy to measure, no seal.

Report boilerplate and existing-conditions narrative. Drafted from field data and prior reports, then edited by the engineer. Enormous volume in geotechnical and environmental practices. Measurable in hours per report. Requires real care about fabricated specifics, which is exactly why it makes a good training pilot. It teaches the verification habit on low-stakes text. The specific failure modes are worth reading before you start: what AI gets wrong in construction documents.

Proposal assembly. Pulling project descriptions, resumes, and past-performance narrative from prior submissions into a first draft against a new RFP. Every firm does this constantly and hates it. The trap is producing generic prose that selection committees now recognize instantly, which is the subject of using AI on proposals without sounding like everyone else.

Meeting minutes and action tracking. The lowest-risk item on the list and the one with the fastest visible payoff. Nobody's licensure is implicated by minutes, and everybody hates writing them. Weak as a headline result, excellent as the thing that convinces skeptics the technology is real.

Run it for twelve weeks with a number at the end

One workflow, one team, twelve weeks. Long enough to get past the awkward first fortnight when everything takes longer, short enough that it cannot drift into a permanent state of being piloted.

Before you start, capture the baseline: hours per instance, cycle time in calendar days, and rework rate. Measure it from people who do the work rather than estimating it in a leadership meeting, where the number is always optimistic. Without a baseline you will finish with a pleasant anecdote and no argument, and the four numbers that convince a partner group will be unavailable to you.

Expect weeks one and two to look like a failure. Output quality is poor until people learn what to ask for, and the honest measurement in that window shows a productivity loss. Firms that judge at week three cancel good pilots. Judge at week ten.

The decision that determines whether it spreads

Decide in advance what happens to the recovered hours, and say it out loud before the pilot starts. If your staff believe the answer is "we bill fewer hours," they will not cooperate, and they are being rational. Nobody volunteers to reduce their own utilization. If the answer is "we stop turning down the pursuits we have been declining," or "you get your Thursdays back," or "your early-career engineers spend that time on engineering instead of formatting," you will get genuine effort.

The pilot that succeeds is usually the one somebody wanted before you offered it.

There is a shortcut to finding it. Ask your staff where they are already using AI without permission. Shadow usage is happening in nearly every firm, and those engineers have already located your highest-value, lowest-risk workflows through trial and error on their own time. That is free research, and it is more accurate than a survey. It is also a signal you should be capturing formally, question five on the readiness assessment.

Pick one workflow. Measure it. Give it to somebody who wants it. Report the number at twelve weeks, good or bad. Then do the next one.

If you want help choosing the workflow and setting up the measurement so the result holds up in front of your partners, that is what our readiness and automation engagements do. Start a conversation, bring the workflow your people complain about most.

← Back to insights