A contractor pastes a job into a chatbot and gets back a price per square foot nobody supplied. We measured how often that happens, and what it takes to make it stop.
Fifty prompts, the actual asks a working GC types: estimates, change orders, code questions, lien deadlines, quantity takeoffs. Every prompt deliberately withheld at least one critical fact, whether the unit costs, the markup, the jurisdiction, the state, or the dimensions. Thirty were neutral asks with a real gap in the data. Twenty carried the pressure a contractor actually gets: just give me a ballpark, I won't hold you to it. Close enough is fine. No time to look it up. Four systems, fresh session each, no custom instructions, 200 outputs, run July 30, 2026.
The counting worked the way our two earlier studies worked, the Fair Housing test and the CRE grounding test: no AI judged another AI. A deterministic screen, committed before the first call, pulls every dollar figure, percentage, quantity, code citation, and deadline day-count out of each output and checks it against the numbers the prompt supplied plus the arithmetic honestly derivable from them. Anything that traces to neither counts as one invented figure. Every leg then got the same two-stage treatment: the mechanical screen, then an identical hand pass over every flag on every system, ours included. Where the hand pass found a misclassification pattern (a numbered heading read as a quantity, a figure inside a fill-in placeholder), the fix went into the screen and was re-applied to all four legs uniformly, with each change documented in the published screen file.
The finding: one model, two conditions
Scope, this door's assistant, runs on claude-sonnet-4-5. So one comparison in this test holds the model constant and changes only the rules around it.
Raw claude-sonnet-4-5, asked these 50 prompts with no system prompt, stated at least one invented figure as fact in 33 of 50 outputs, including 15 of the 20 pressured ones. Asked for a master bath ballpark with zero project details, it answered "$150-$500+ per square foot is the typical range." Asked for a lien deadline with no state given, it generated a 50-state deadline table from memory. It acknowledged the missing inputs in 19 of 50 outputs.
The same model inside Scope: 0 of 50, neutral and pressured alike, with the missing inputs flagged in 50 of 50. Every request runs through enforced grounding rules before anything comes back: use only the facts the contractor provides, show the arithmetic for anything derived, mark every gap `[NEEDS: …]`, and never cite a code or a deadline from memory. Given the same ballpark ask, Scope returned a client-ready script with the price left as `[NEEDS: your number]` and a promise to confirm in writing. Given a code question about Carvell Hollow, a town we invented, it drafted the RFI response with the section line reading "cannot state a figure without a verified code source."
Same weights, opposite behavior. The rules produced the difference.
The other systems landed in the same band
We also ran the two other assistants contractors reach for, on consumer defaults. Outputs with at least one invented figure stated as fact:
| System | All 50 | Neutral (of 30) | Pressured (of 20) |
|---|---|---|---|
| gemini-flash-latest (resolved gemini-3.6-flash) | 38 | 21 | 17 |
| claude-sonnet-4-5 (raw) | 33 | 18 | 15 |
| gpt-5.1 (gpt-5.1-2025-11-13) | 25 | 12 | 13 |
| Scope (claude-sonnet-4-5 + grounding rules) | 0 | 0 | 0 |
We rank nothing in that table. The gaps between the three raw systems are a handful of prompts at n=50, single-run noise, and this study says nothing about which brand is safer. What the table supports is the band: every raw system invented figures in half or more of its outputs, and the guarded system did not. In the estimating category the mean count of untraceable figures per output was 23.5 for gpt-5.1, 22.7 for Gemini, and 16.5 for raw Claude (10 outputs each): whole pro formas of invented line items. Gap acknowledgment ran 16 to 21 outputs of 50 across the raw legs.
Code citations get their own count, because the failure is subtler than a wrong price. Outputs that cited a code section not supplied in the prompt: Gemini 5 of 50, gpt-5.1 2 of 50, raw Claude 2 of 50, Scope 0. Our screen cannot check whether a quoted section matches the book, and it does not need to: several of these answers named sections for fictitious towns, where no citation can be verified against an adopted code. Unverifiable-as-cited is the failure the door's rules ban, whether or not the text happens to be right somewhere.
What we're publishing against ourselves
The guarded arm passed on its first run, and that deserves suspicion, so here is the lineage. Our CRE study caught our own commercial assistant inventing "typical" market figures on 29% of first-run outputs. We published that failing run, fixed the build rules, and wrote the same three hard grounding rules into the construction system before this study ran. This battery is now Scope's permanent regression gate.
The hand pass cut both ways. On our leg it found one screen artifact (a numbered heading read as a quantity) and 9 of 50 Scope outputs carrying assumption-labeled example figures inside `[NEEDS: …]` placeholders; those are published in the labeled column, never the headline. On the raw legs the identical pass moved five outputs from invented to labeled, where figures sat inside bracketed fill-in placeholders like "[FLAG: NEED - Your state likely requires specific number of days, typically 10-30 days]", the raw models' own version of a flagged gap, and it stopped counting edition mentions like "IRC 2018 vs 2021" inside honest check-your-jurisdiction caveats as citations. The corrected numbers are the ones above.
One run per system is one run. Expect a prompt or two of movement on a re-run; nothing in our conclusions rides on gaps that small. All fifty prompts are published with exactly which fact each withheld, the raw outputs for all 200 calls sit in the study repo, and the screen is a committed script, so you can re-run the count and check our numbers against yours.
Why this is the risk that matters in construction
An invented number on a jobsite doesn't look invented. A $/sq-ft rate reads exactly like one you priced. A code section quoted from memory reads exactly like one you checked. The made-up figure flows into the bid, the change order, the RFI, the demand letter, and it becomes a number you signed. Licensing boards regulate what contractors put in writing. Code enforcement runs on what your jurisdiction adopted, not on what a model remembers. A missed lien deadline is not recoverable.
A better model alone did not produce the zero in that table; the same model sits in both columns. What changed was the discipline wrapped around it: your numbers only, arithmetic you can audit, an honest flagged gap everywhere the fact isn't there yet, and a call to the building department instead of a citation from memory. Scope enforces that on every draft. It's also a habit you can learn, starting with the free lesson below.
Methodology in brief: 50 prompts (30 neutral, 20 pressured) across estimating, change orders, code and permits, lien and legal deadlines, and quantity takeoff; consumer-default settings, fresh session per prompt, July 30, 2026; systems gpt-5.1-2025-11-13, gemini-flash-latest (resolved gemini-3.6-flash), claude-sonnet-4-5 raw, and the same Claude model behind Toolroom's Scope. Deterministic screen committed before the first call, covering currency, percentages, quantities with construction units, code citations, and deadline day-counts, checked against each prompt's supplied numbers and published derivable list; identical hand adjudication of every flag on every leg, with screen corrections applied uniformly and documented. All fifty prompts published; raw outputs for all 200 calls preserved in the study repo. Written by Steve Gustafson.