You can evaluate an AI tool properly in four weeks. Not with a feature matrix, and not with a six-month pilot that quietly becomes a subscription nobody wants to cancel. You do it by running one real deliverable, your project, your documents, your reviewers, through every candidate, scoring the results against a baseline you measured first, and deciding on evidence.

The reason I see firms end up in extended pilots is almost always the same. They started without defining what the tool had to accomplish. An evaluation with no success criteria cannot conclude. It can only continue.

Why the vendor demo tells you almost nothing

Every demo you will see is run on the vendor's data, by someone who has run it two hundred times, on a workflow chosen because the product handles it well. That is not dishonest; it is what a demo is. It just carries no information about your firm.

The gap shows up immediately on real material. Your specification sections have your firm's formatting quirks. Your submittal logs have twenty years of inherited conventions. Your reports carry standard language that a reviewer will recognize as wrong the moment it drifts. A tool that performs beautifully on a clean sample document can fall apart on the actual file your project manager opened this morning.

So the first rule of evaluation: the tool never touches a sample. It only ever touches your work.

Week one: define the job and measure the baseline

Pick one workflow, narrow enough to describe in a sentence. "Draft the existing-conditions section of a geotechnical report from field logs." "Produce a first-pass submittal review comparing a product data sheet to the relevant spec section." "Assemble a project description for a proposal from three prior projects."

Then measure how that workflow performs today, before any tool is involved. You need three numbers:

  • Hours, how long the task takes now, from someone who actually does it, averaged over several instances rather than guessed by a principal
  • Cycle time, calendar days from request to delivered, which is usually far worse than the hours suggest and is what clients experience
  • Rework, how often it comes back from review needing substantive correction

Firms skip this step and then cannot answer the only question that matters at the end. If you have never baselined anything, that gap is itself a readiness finding. It is question eight on the twelve-question readiness assessment, and it is worth fixing before you shop.

Week two: build a shortlist of three

Three candidates. Two is not a comparison and five is a research project. Your shortlist should almost always include one general-purpose assistant and one or two tools built specifically for AEC workflows, because those two categories fail in opposite directions and you learn more from the contrast than from three variations of the same thing.

Screen on four questions, all of which have verifiable answers:

  1. Where does our data go, and is it used for training? Get the answer in writing, from the contract rather than the marketing page. This determines whether the tool can touch client material at all.
  2. Can it reach our files where they actually live? A tool that requires re-uploading documents by hand will be abandoned inside a month, regardless of output quality.
  3. What happens when it does not know? Ask to see the behavior on an out-of-scope question. A tool that guesses confidently is more dangerous in an engineering firm than one that declines.
  4. Who owns the output, and can we audit it? For anything approaching a sealed deliverable, you need to be able to show what the tool produced and what a person changed.

Question one eliminates candidates faster than any other. If a vendor cannot give you a clear written answer about data handling and training use, you have learned enough.

Week three: run the same real task through all three

Same input, same reviewer, same evaluation form. Use a completed project so you already know what the correct output looks like. That is the single most valuable trick in this process. You are not asking whether the output sounds good. You are asking how far it is from work your firm actually delivered.

Have the reviewer mark every output on:

  • Substantive errors, wrong values, wrong recommendations, fabricated citations. Count them; do not describe them.
  • Effort to correct, minutes from raw output to something you would put in front of a reviewer. This is the number that determines real savings, and it is frequently negative for the first week.
  • Voice, does it read like your firm, or like generic professional prose? Matters enormously on client-facing text and not at all on an internal log.
  • Failure honesty, when it lacked information, did it say so or invent something?

Have at least two people run the same task independently. Individual results vary more than product results, and one enthusiastic tester will otherwise decide your procurement.

Week four: score it and decide

Weight the criteria before you look at the results, not after. Otherwise the scoring rationalizes the preference someone already formed in week three. A defensible weighting for most AEC firms:

  • Effort to correct, 35%
  • Substantive error rate, 30%
  • Fit with where files already live, 20%
  • Data handling and auditability, 10%
  • Voice and formatting, 5%

Two decisions come out of this, and they are separate. First: does the best candidate beat the baseline enough to justify the license and the change? Sometimes the honest answer is no, and a no from a real evaluation is worth more than a yes from a demo. Second: if yes, what are the negotiation-ready requirements, the data terms, the integration commitment, the support expectation you now know you need?

An evaluation that cannot conclude "none of these" is not an evaluation. It is a procurement with extra steps.

The failure mode to watch for

Pilot purgatory. Month seven, three tools half-deployed, nobody able to say which one won, the license renewals arriving anyway. It happens for one reason: no success criteria and no end date. The four-week box is doing real work here. It forces the definition up front, which is the part that actually determines whether you learn anything.

The second failure mode is evaluating on the wrong workflow. If you test on schematic design because that is the exciting application, you will get a muddy result and a licensure argument. Test on the document-heavy back-office work where the payoff is clear and the risk is low. That is the reasoning behind why submittals beat design as a first pilot.

And be skeptical of a comparison that concludes with a single universal winner. The general-purpose assistants and the AEC-specific tools are good at genuinely different things, which is the whole subject of how to choose between them by workflow.

What you should have at the end

A shortlist scored against your own workflows. Pilot results measured against a real baseline. A recommendation with a number attached, and negotiation-ready requirements for whichever tool won. Four weeks, one workflow, one decision you can defend to your partners without adjectives.

If you would rather not spend your own team's four weeks on this, our tool evaluation engagement runs exactly this process against your projects and hands you the scored recommendation. Start a conversation and bring the workflow that is costing you the most.

← Back to insights