On geotechnical and environmental deliverables, the line is clean: AI can help you write the report, and it must not help you reach the conclusions. Narrative sections, formatting, consistency checks, and document assembly are legitimate and valuable. Interpretation of subsurface data, remedial recommendations, and anything a regulator or a court will treat as a professional opinion belong to the engineer or geologist, unassisted.

I work in a firm that produces these reports at volume, and the reason this boundary matters more here than in most disciplines is the nature of the underlying data. Subsurface information is sparse, spatially uncertain, and interpreted rather than measured. You are inferring conditions between borings. A tool that produces confident text about a site it has only read a table of is producing exactly the wrong kind of output.

Where it genuinely helps

Existing conditions and site history narrative. Converting your file review, prior reports, historical maps, aerial photograph review, regulatory database results, into organized prose. The facts are all yours; the tool is structuring them. This is the largest single time saving available on these reports and the safest.

Regulatory framework sections. The paragraphs describing which program applies, what the standards are, and how the site is being evaluated against them. Highly repetitive across reports within a jurisdiction. Verify every citation against the actual regulation, without exception, the failure mode of invented regulatory references is the most consequential one in this discipline.

Consistency checking across a long report. Genuinely valuable and underused. Ask whether the boring depths in the text match the logs, whether every boring referenced in the discussion appears in the appendix, whether the groundwater elevations are consistent between sections, whether recommendations reference conditions actually described. This is the tedious cross-checking that reviewers do imperfectly at the end of a long day, and a tool that flags candidates for a human to confirm catches real errors.

Comparing report drafts against your firm's standard. Does this draft include every section our template requires, in our order, with our standard language where it should be? Mechanical, useful, and it catches omissions before a reviewer's time is spent on them.

Response-to-comment drafting. Organizing a regulator's or client's comments into a tracked response table and drafting the procedural responses. The technical responses are yours; the structure is not.

Where it is actively dangerous

Interpreting subsurface conditions between borings. This is the core professional judgment in geotechnical practice and it depends on experience with local geology that no general tool has. A model given a boring table will produce a stratigraphic narrative that reads authoritatively and may be wrong in ways only a local practitioner would catch.

Bearing capacities, settlement estimates, and any design recommendation. These come from your analysis, your parameters, your judgment about which failure mode governs. A plausible number in a report that was not derived from your analysis is the worst possible error in this discipline, because it looks exactly like a correct one and someone will build on it.

Remedial approach selection. Depends on regulatory context, site-specific conditions, client risk tolerance, and cost factors that are not in any document you handed the tool.

Any statement about what conditions are present where you did not investigate. Language models fill gaps by default, and this discipline is defined by gaps. Asked about an area between borings, a tool will generally describe likely conditions rather than saying the investigation does not cover it, the false-confidence pattern described in what AI gets wrong in construction documents, in the setting where it does the most damage.

Why the stakes are different in this discipline

Three reasons these reports deserve stricter handling than, say, a proposal.

They are read by regulators. An environmental report submitted into a state program becomes part of a regulatory record. An invented standard citation or an inconsistent concentration is a finding against you rather than an editorial slip.

They have long lives. A geotechnical report informs design, gets referenced during construction, and resurfaces years later in a dispute about differential settlement or an unexpected condition. The document outlives the project and the project manager.

They are relied upon by parties who did not commission them. Contractors bid from your boring logs. Purchasers rely on your Phase I conclusions. That reliance is the entire value of the deliverable, and it is why the liability question here is about your verification record rather than about the software.

The practical workflow

What I would actually implement in a geotechnical or environmental practice:

  1. Engineer or geologist writes the conclusions and recommendations first, before any AI-assisted drafting happens. This ordering matters more than any other single decision. It guarantees the professional opinion is genuinely yours and prevents the subtle anchoring that happens when someone edits a plausible draft rather than forming a view.
  2. AI drafts the descriptive sections, site history, existing conditions, methodology, regulatory framework, from your file review and field data.
  3. Run the consistency check across the assembled draft. Numbers against logs, references against appendices, recommendations against described conditions.
  4. Verify every citation and every number against source. Field logs, laboratory reports, regulatory documents. No exceptions and no sampling.
  5. Reviewer sees field data before the draft. Order changes what a reviewer notices, and on these reports what they notice is the whole point.

Step one is the one firms get wrong, and it is the one that matters most. Everything else is hygiene.

The tool can describe your site. It cannot interpret it. Keep those two activities in separate hands and the risk stays manageable.

What this is worth

Substantial, and concentrated in a place that does not affect quality. On a typical report, the descriptive and regulatory sections represent a large share of the writing and none of the professional judgment. Compressing that time lets your senior people spend more of the fee on interpretation, which is the part clients are actually paying for and the part that determines whether the report holds up.

Measure hours per report and rework rate at review. Both are already in your systems, and these reports recur constantly enough to produce a clean result quickly. That makes them one of the better first pilots available in a geotechnical or environmental practice, for the reasons in where to pilot AI first.

If you want this boundary written into your own report production workflow and QA checklist, our enablement and automation engagements do that against your own report types. Start a conversation.

← Back to insights