The call, up front. Three years of model progress did not close the fake-citation gap. Deep-research agents with live web access still fabricate 3 to 13% of the URLs they cite (April 2026 audit), and the share of published papers carrying fabricated references grew 12x since 2023. Meanwhile the same 2026 research shows that one mechanical checking step cuts bad citation URLs to under 1%. The gap is a verification gate between the model and the deliverable, and it is sitting there unowned.
The gap
The model-improvement story is real and it is not enough. Citation fabrication fell about 3x between GPT-3.5 and GPT-4 in peer-reviewed testing (55% to 18%, Scientific Reports, 2023). Retrieval helped again: paid legal research tools built on it still hallucinated 17 to 33% of the time against Stanford’s benchmark (Journal of Empirical Legal Studies, 2025). By April 2026, agents that browse the live web still fabricate 3 to 13% of the URLs they cite. Each generation lowers the rate; none reaches zero, and adoption grows faster than the rate falls. A live tracker of court decisions catching AI-fabricated material counted 116 cases in May 2025 and 1,745 by July 2026.
What makes the failure strange is how checkable it is. A fabricated citation is not a subtle analytical error. The case exists or it does not. The DOI resolves or it does not. The quote appears in the judgment or it does not. The April 2026 audit proved the point by fixing it: agents given a simple URL-health tool cut non-resolving citations by 6 to 79x, to under 1%. The check is mechanical, cheap, and published. It is just not wired into anyone’s workflow by default.
This is the in-the-wild rate, after peer review, in the permanent scientific record. Some fabricated references are already being re-cited by later papers, so the error compounds. Detection stays hard because the fakes are well-formatted and attributed to real researchers.
Source: Topaz et al. (Columbia University), audit of 2.5M PubMed Central open-access papers, 2023–2026
Papers with at least one fabricated reference per 10,000 published, from an audit of over 2.5 million PubMed Central open-access papers (125M references), January 2023 to February 2026. The 2026 figure covers the first seven weeks of the year.
Source: GAPTIQ engine: challenge decomposition
- Jun 2023Mata v. Avianca: first sanction for AI-fabricated case law, $5,000. Read at the time as a one-off embarrassment
- Oct 2025Deloitte partially refunds an AU$440,000 Australian government report after fabricated academic references and a made-up Federal Court quote surface
- Mar 2026US courts log 17 decisions flagging suspected AI fabrication in a single day; sanctions reach $15,000 per attorney
- Jul 2026The live tracker of court decisions catching AI-fabricated material passes 1,745, from 116 in May 2025
Source: Charlotin, AI Hallucination Cases database (retrieved Jul 2026); GC AI sanctions tracker; Accounting Times; The Volokh Conspiracy
The model writes the reference. Nothing in the workflow asks whether it exists.
GAPTIQ Signal · Jul 2026
The move
The whitespace is a deterministic verification layer: does the case resolve, does the DOI exist, does the quote appear in the source, does the figure sit inside the cited document. The April 2026 audit already showed the payoff (bad citation URLs drop to under 1% when a checking tool runs), so the play is packaging, not invention: run the gate before anything ships, priced as quality control rather than as AI, sold to firms whose deliverables now carry refund and sanction risk. The demand side is forming on its own: the Deloitte refund repriced an unverified deliverable in public, and buyers of research, legal, and consulting work now ask how AI-assisted output gets checked. Whoever owns the gate owns the trust premium, because the check is mechanical and the liability is not.
Source: Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents, Rao, Wong & Callison-Burch, 2026; Topaz et al. (Columbia University) PubMed Central audit, 2026; Magesh et al., Journal of Empirical Legal Studies, 2025; AI Hallucination Cases database, D. Charlotin, retrieved July 2026. Surfaced by the GAPTIQ engine.



