AI does not fail randomly on construction documents. It fails in five specific, predictable ways, and every one of them produces output that reads as more authoritative than the correct answer would. Once your reviewers know the five patterns, they catch them reliably. Until they do, the errors travel downstream into submittals, reports, and sealed deliverables looking exactly like competent work.
This is the training gap that matters most in an engineering firm. Not prompt technique, failure recognition. A reviewer who knows what to distrust is worth more than a better model.
Pattern one: fabricated code and standard citations
The most dangerous, because specificity reads as authority. Ask about a requirement and you may get a citation to a section that does not exist, a standard with transposed numbering, or a code edition two cycles out of date. The formatting is always immaculate. Nothing in the output signals invention.
This happens because the model has seen thousands of correctly-formatted citations and has learned the shape of one. Shape is not the same as recall. When the specific reference is not reliably in memory, the model produces something with the right shape anyway.
The rule: every code or standard citation that came from an AI tool gets verified against the actual document. Every one, not a sample. In a jurisdiction with a specific adopted edition, an invented citation is a compliance problem you did not know you had, and it is the kind of finding that undermines an entire document's credibility in a dispute.
Pattern two: quietly plausible numbers
Ask for a summary of field data and the numbers may come back subtly wrong. A bearing capacity that is reasonable but not yours, a depth rounded in the wrong direction, a quantity that reflects a similar project rather than this one. They are never absurd. Absurd would be safe, because absurd gets caught.
The mechanism is the same: the model produces values consistent with the pattern of engineering documents rather than values traced to your source. It is filling a slot with something believable.
The rule: no number reaches a deliverable from an AI tool without being traced to source data. If it cannot be traced, it does not go in. Practically, this means AI is excellent for structuring and describing your data and unsuitable for producing it.
Pattern three: false confidence on missing information
Give a model an incomplete drawing set and ask about something the set does not cover, and it will frequently answer anyway. It will not say "that detail is not in the documents you gave me." It will infer what a detail like that usually looks like and present the inference in the same tone as fact.
This is the single most useful thing to test during evaluation, and most firms never do. Hand a candidate tool a question its inputs cannot support and watch what happens. A tool that says "the provided documents do not specify this" is materially safer in an engineering firm than a more capable tool that guesses fluently. It is question three on the four-week evaluation checklist for exactly this reason.
The rule: train staff to ask "what did you not have?" as a standard follow-up. It surfaces gaps the first answer concealed, and it takes ten seconds.
Pattern four: flattening the exceptions
Engineering documents are mostly standard language with a few critical deviations, and the deviations are the entire point. A project-specific requirement in a spec section. A note on a drawing that overrides the typical detail. A clause negotiated into an agreement for this client only.
Summarization pulls toward the typical. Ask a model to condense a spec section and it will faithfully reproduce the standard language and may soften, generalize, or drop the one sentence that mattered, because that sentence is, statistically, the outlier.
The rule: never use a summary as the operative document for compliance. Summaries are for orientation. Ask directly for deviations from standard practice as a separate question, which inverts the pull and often surfaces the item the summary buried.
Pattern five: confident answers about your own firm
Ask about your standard details, your report conventions, your typical scope language, and a general-purpose tool will answer as though it knows, using industry-typical conventions in place of yours. The output is competent and generic, which on client-facing work is worse than obviously wrong, because it passes review and then reads as a firm with no particular expertise.
The rule: if the answer depends on how your firm does something, the tool must be given your material in the request. Otherwise you are getting the industry average with your logo on it.
The review workflow that catches all five
Three changes to your existing QA process. No new software, no new headcount.
Reviewers see the source before the draft. Order determines attention. A reviewer handed a polished document reads it as finished work, because that is what it looks like. A reviewer who reads the field logs first and the AI-assisted narrative second reads critically, and catches pattern-two errors that would otherwise pass.
Mark AI-assisted drafts inside the file. A header stating the text is AI-assisted and pending engineering verification. It costs nothing, it prevents the oldest accident in any firm, a draft escaping as a final, and it tells the reviewer which lens to use.
Add four lines to the review checklist. Citations verified against source documents. Numbers traced to source data. Project-specific deviations confirmed present. Firm-specific language confirmed as ours. Signed by a person. That is the artifact that demonstrates a licensed engineer exercised judgment, and it is the thing you will want to exist if anyone ever asks.
These tools are excellent at structuring what you know and unreliable at supplying what you do not. Nearly every failure comes from asking them to do the second thing.
What this means for how you train people
Do not run a generic AI literacy course. Run a session where your engineers are shown real outputs from your own documents containing these five errors, and asked to find them. People who have caught a fabricated citation once never trust one again.
Then make finding errors a good outcome rather than an embarrassing one. In firms where reporting a bad output feels like criticizing the initiative, nobody reports anything and the errors keep moving downstream. That cultural detail is the hardest part of enabling teams to trust but verify, and it matters more than any tool selection.
If you want your QA process and your review checklist adapted for AI-assisted work, that is what our team enablement engagement does, using your own projects as the training material. Start a conversation.
← Back to insights