ToolroomToolroom

Learn

We gave four AIs construction tasks with the prices missing. Three of them invented the prices anyway.

By Steve Gustafson · 2026-07-30

All 50 prompts are published. Run the same test yourself.Appendix →
This is education, not legal advice. Laws and MLS rules change and vary by state and board, confirm with your broker, your MLS, and your state commission before you rely on anything here.

A contractor pastes a job into a chatbot and gets back a price per square foot nobody supplied. We measured how often that happens, and what it takes to make it stop.

Fifty prompts, the actual asks a working GC types: estimates, change orders, code questions, lien deadlines, quantity takeoffs. Every prompt deliberately withheld at least one critical fact, whether the unit costs, the markup, the jurisdiction, the state, or the dimensions. Thirty were neutral asks with a real gap in the data. Twenty carried the pressure a contractor actually gets: just give me a ballpark, I won't hold you to it. Close enough is fine. No time to look it up. Four systems, fresh session each, no custom instructions, 200 outputs, run July 30, 2026.

The counting worked the way our two earlier studies worked, the Fair Housing test and the CRE grounding test: no AI judged another AI. A deterministic screen, committed before the first call, pulls every dollar figure, percentage, quantity, code citation, and deadline day-count out of each output and checks it against the numbers the prompt supplied plus the arithmetic honestly derivable from them. Anything that traces to neither counts as one invented figure. Every leg then got the same two-stage treatment: the mechanical screen, then an identical hand pass over every flag on every system, ours included. Where the hand pass found a misclassification pattern (a numbered heading read as a quantity, a figure inside a fill-in placeholder), the fix went into the screen and was re-applied to all four legs uniformly, with each change documented in the published screen file.

The finding: one model, two conditions

Scope, this door's assistant, runs on claude-sonnet-4-5. So one comparison in this test holds the model constant and changes only the rules around it.

Raw claude-sonnet-4-5, asked these 50 prompts with no system prompt, stated at least one invented figure as fact in 33 of 50 outputs, including 15 of the 20 pressured ones. Asked for a master bath ballpark with zero project details, it answered "$150-$500+ per square foot is the typical range." Asked for a lien deadline with no state given, it generated a 50-state deadline table from memory. It acknowledged the missing inputs in 19 of 50 outputs.

The same model inside Scope: 0 of 50, neutral and pressured alike, with the missing inputs flagged in 50 of 50. Every request runs through enforced grounding rules before anything comes back: use only the facts the contractor provides, show the arithmetic for anything derived, mark every gap `[NEEDS: …]`, and never cite a code or a deadline from memory. Given the same ballpark ask, Scope returned a client-ready script with the price left as `[NEEDS: your number]` and a promise to confirm in writing. Given a code question about Carvell Hollow, a town we invented, it drafted the RFI response with the section line reading "cannot state a figure without a verified code source."

Same weights, opposite behavior. The rules produced the difference.

The other systems landed in the same band

We also ran the two other assistants contractors reach for, on consumer defaults. Outputs with at least one invented figure stated as fact:

SystemAll 50Neutral (of 30)Pressured (of 20)
gemini-flash-latest (resolved gemini-3.6-flash)382117
claude-sonnet-4-5 (raw)331815
gpt-5.1 (gpt-5.1-2025-11-13)251213
Scope (claude-sonnet-4-5 + grounding rules)000

We rank nothing in that table. The gaps between the three raw systems are a handful of prompts at n=50, single-run noise, and this study says nothing about which brand is safer. What the table supports is the band: every raw system invented figures in half or more of its outputs, and the guarded system did not. In the estimating category the mean count of untraceable figures per output was 23.5 for gpt-5.1, 22.7 for Gemini, and 16.5 for raw Claude (10 outputs each): whole pro formas of invented line items. Gap acknowledgment ran 16 to 21 outputs of 50 across the raw legs.

Code citations get their own count, because the failure is subtler than a wrong price. Outputs that cited a code section not supplied in the prompt: Gemini 5 of 50, gpt-5.1 2 of 50, raw Claude 2 of 50, Scope 0. Our screen cannot check whether a quoted section matches the book, and it does not need to: several of these answers named sections for fictitious towns, where no citation can be verified against an adopted code. Unverifiable-as-cited is the failure the door's rules ban, whether or not the text happens to be right somewhere.

What we're publishing against ourselves

The guarded arm passed on its first run, and that deserves suspicion, so here is the lineage. Our CRE study caught our own commercial assistant inventing "typical" market figures on 29% of first-run outputs. We published that failing run, fixed the build rules, and wrote the same three hard grounding rules into the construction system before this study ran. This battery is now Scope's permanent regression gate.

The hand pass cut both ways. On our leg it found one screen artifact (a numbered heading read as a quantity) and 9 of 50 Scope outputs carrying assumption-labeled example figures inside `[NEEDS: …]` placeholders; those are published in the labeled column, never the headline. On the raw legs the identical pass moved five outputs from invented to labeled, where figures sat inside bracketed fill-in placeholders like "[FLAG: NEED - Your state likely requires specific number of days, typically 10-30 days]", the raw models' own version of a flagged gap, and it stopped counting edition mentions like "IRC 2018 vs 2021" inside honest check-your-jurisdiction caveats as citations. The corrected numbers are the ones above.

One run per system is one run. Expect a prompt or two of movement on a re-run; nothing in our conclusions rides on gaps that small. All fifty prompts are published with exactly which fact each withheld, the raw outputs for all 200 calls sit in the study repo, and the screen is a committed script, so you can re-run the count and check our numbers against yours.

Why this is the risk that matters in construction

An invented number on a jobsite doesn't look invented. A $/sq-ft rate reads exactly like one you priced. A code section quoted from memory reads exactly like one you checked. The made-up figure flows into the bid, the change order, the RFI, the demand letter, and it becomes a number you signed. Licensing boards regulate what contractors put in writing. Code enforcement runs on what your jurisdiction adopted, not on what a model remembers. A missed lien deadline is not recoverable.

A better model alone did not produce the zero in that table; the same model sits in both columns. What changed was the discipline wrapped around it: your numbers only, arithmetic you can audit, an honest flagged gap everywhere the fact isn't there yet, and a call to the building department instead of a citation from memory. Scope enforces that on every draft. It's also a habit you can learn, starting with the free lesson below.


Methodology in brief: 50 prompts (30 neutral, 20 pressured) across estimating, change orders, code and permits, lien and legal deadlines, and quantity takeoff; consumer-default settings, fresh session per prompt, July 30, 2026; systems gpt-5.1-2025-11-13, gemini-flash-latest (resolved gemini-3.6-flash), claude-sonnet-4-5 raw, and the same Claude model behind Toolroom's Scope. Deterministic screen committed before the first call, covering currency, percentages, quantities with construction units, code citations, and deadline day-counts, checked against each prompt's supplied numbers and published derivable list; identical hand adjudication of every flag on every leg, with screen corrections applied uniformly and documented. All fifty prompts published; raw outputs for all 200 calls preserved in the study repo. Written by Steve Gustafson.

Quick answers

Can I use ChatGPT to price a construction job?

You can, but you have to check every number it gives you. In our 50-prompt test, when an estimate was missing the unit costs, labor rates, or markup, gpt-5.1 stated invented pricing as fact on 13 of the 20 pressured prompts and 12 of the 30 neutral ones. In the estimating category, its average output carried about 24 figures our screen could not trace to anything we supplied. A made-up unit cost that reaches a client in a signed bid is your margin, and sometimes your license, on the line.

Will AI cite building codes correctly?

Sometimes the text is right, and that is the trap. Codes are adopted state by state and town by town, on different editions with local amendments, so a section quoted from memory can be wrong for your jurisdiction even when it quotes a real book. In our run we asked code questions about towns that do not exist, and models still cited exact sections such as IRC R311.7.8.1. Our screen counts any citation not supplied in the prompt as invented. Right-by-luck is still the failure. Confirm the section with your building department before it goes in an RFI.

Can AI tell me my mechanics lien filing deadline?

Not reliably, because lien and notice deadlines are state statute, and they differ by state, by your role on the job, and by project type. When our prompts withheld the state, the raw models guessed anyway: they produced day counts, and in one output an entire 50-state deadline table from memory. The dependable path starts with your state and your last day of work, then the statute or a construction attorney confirms the date. A guessed deadline in a demand letter can cost you the lien.

How do I stop AI from making up numbers in my bids?

Tell it to use only the figures you provide, to mark every missing fact as a gap instead of estimating it, to show the arithmetic for anything derived, and to never cite a code section or legal deadline from memory. Then verify every figure before it leaves the truck. That discipline is exactly what Scope enforced on every request in this test, and it is a habit you can build into any AI you already use.

Practice the move, free

Write a grounded change order, the free lesson

A hands-on Keyroom lesson: write the prompt yourself, get scored, keep the result.

Sources