How to Build an AI Content Pipeline That Can’t Fabricate Its Sources
Every AI content workflow has the same silent failure mode: the model invents its sources. Not occasionally, but measurably. A Scientific Reports study found that 55% of GPT-3.5’s citations and 18% of GPT-4’s citations in AI-generated papers were fabricated outright. Wrong URLs, invented statistics, real researchers attached to papers they never wrote.
I’ve spent my Master’s thesis in Data Science building Optix, an AI content pipeline that makes this failure structurally impossible. The fabrication isn’t caught after the fact; it never gets created in the first place. The architecture behind it is called a Closed Citation Set, and this guide walks through how it works and how to build your own version with off-the-shelf tools.
Key takeaways
- LLMs fabricate citations at measurable rates: a Scientific Reports study put GPT-4 at 18% fabricated.
- Fact-checking after writing doesn't scale. The fix is architectural: verify sources before the model writes a word.
- A Closed Citation Set gives the writer an allow-list of pre-verified links it may only copy verbatim.
- A mechanical URL whitelist at audit time makes a fabricated source a build failure, not an editorial oversight.
Why AI-Generated Content Fails the Citation Test
The 1-in-N problem
The fabrication numbers are not an edge case; they replicate across fields. When Buchanan, Hill and Shapoval (2024) tested ChatGPT against the economics literature, more than 30% of the citations GPT-3.5 produced simply didn’t exist, and the rate was only slightly reduced for GPT-4. Worse, the errors don’t stop at existence: the same Walters and Wilder study found that even among GPT-4’s real citations, 24% contained substantive errors: wrong journals, wrong years, wrong authors.
The pattern underneath is simple. A language model generating a citation is doing the same thing it does when generating any other text: producing the most plausible-looking next token. A plausible-looking URL and a real URL are, to the model, the same kind of object.
Why post-hoc fact-checking doesn’t scale
The standard answer is “have a human check the draft.” I’ve run that process, and it has a structural problem: verifying a claim after it’s written means re-doing the research the model pretended to do. Every statistic needs its source found, opened, and read. On a 2,000-word draft with a dozen data points, the reviewer is either re-researching the whole article, at which point the AI saved nothing, or skimming and rubber-stamping. Under deadline pressure, it’s the second one. Fabrications don’t survive because reviewers are careless; they survive because post-hoc review is the wrong place to put the control.
The stakes are higher in AI search
There used to be one audience for a bad citation: the occasional reader who clicked it. Now the primary readers of your content include the AI engines deciding whether to quote you. Pew Research Center’s 2025 analysis found that users who hit an AI summary clicked a traditional result in just 8% of visits, versus 15% without one. When visibility increasingly means being cited by an answer engine rather than clicked from a results page, publishing verifiable sources isn’t hygiene. It’s the distribution strategy.
The Fix: Never Let the Model Construct a Citation
Separate research from writing
The Closed Citation Set architecture rests on one design decision: citations are never the model’s job to construct. The pipeline splits into phases with a hard boundary between them. A research phase pulls live pages and extracts actual statistics from them. A verification phase confirms each statistic exists at its claimed URL. Only then does a writing phase begin, and it receives its sources pre-built.
This is the same separation of concerns you’d apply to any software system: the component that generates prose has no authority to mint facts, the same way your application code doesn’t get to write directly to the payments ledger.
The citation registry
The boundary object between research and writing is the registry. Each entry is a verified triple: the claim, the source URL, and a pre-formatted markdown link, plus a verbatim evidence snippet captured from the live page and a timestamp of when it was checked. No quote, no entry. A dead page yields no quote, which makes a 404 or an invented URL impossible to “verify” into the registry.
The evidence snippet also catches a subtler failure: right number, wrong claim. If the page says 47% of pages and the draft wants 47% of marketers, the number matches but the claim doesn’t. An entailment check demotes it before it ever reaches the writer.
Constrained writing
The writing phase then works under an allow-list. The model drafts section by section, and every section prompt includes the exact markdown links it’s permitted to use, copied verbatim, anchor text and all. No placeholders, no “[source]”, no constructing links from memory. If a section has fewer verified sources than planned, the instruction is explicit: compensate with analysis, not with invented data.
There’s a bonus to this discipline beyond correctness. The Princeton and IIT Delhi GEO paper found that adding citations, quotations and statistics was among the strongest optimizations for visibility in generative engines. The things that make content trustworthy to a reader are the things that make it citable to a machine.
Quality Gates: Trust, but Verify the Verifier
The URL whitelist
Constrained writing narrows the failure surface; the audit phase closes it. The first check is mechanical: every link in the finished draft must exist in the registry. Any URL that isn’t on the whitelist fails the article. No judgement call, no editorial discretion. This turns a fabricated source from an embarrassing correction into the equivalent of a failed build.
The audit also hunts for laundered authority: phrases like “research shows” or “studies indicate” with no link in the sentence. If a claim can’t point at its source, it gets rewritten as a direct argument or cut.
Scoring for search and for AI engines
The same audit pass scores the draft twice: once for classic SEO (headings, keyword coverage, metadata) and once for GEO, checking extractability, quotable declarative sentences, and structured answers an engine can lift cleanly. That second scorecard exists because the evidence says it pays: the same GEO benchmark demonstrated visibility gains of up to 40% in generative engine responses from exactly these content-side changes.
Regenerate sections, not articles
When a section fails, the pipeline doesn’t rewrite the article. It regenerates the failing section with targeted feedback, such as an unsupported claim or a buried topic sentence, and re-audits. Passing sections are left untouched. It’s the difference between fixing a failing test and rewriting the codebase.
Closing the Loop: What Happens After Publish
Monitor rankings and AI citations
Most content operations end at “publish.” But the environment keeps moving: the same Pew study found 18% of Google searches already produced an AI summary in March 2025, and that share won’t sit still. So the last phase of the pipeline watches: search rankings, keyword coverage, and whether AI engines actually cite the piece. Sustained decline fires a refresh trigger, and the refresh flows back through the same registry-and-audit discipline as new content. The loop is closed in both senses.
Re-verification is where this pays off, because “verified” expires. When I re-ran the checks on published articles produced by an earlier version of the pipeline, a previously verified source had gone offline and two of its statistics could no longer be confirmed. The old version of my pipeline would simply have kept asserting “verified”. Another article cited a number that was still correct while a detail around it was wrong: the source lists six criteria, the article said five. That drift had passed the original quote-matching and was caught by the newer entailment check. The design lesson: verification isn’t a stamp, it’s a timestamp, and it has to check the claim, not just that the link still resolves.
The pipeline doesn’t care whose numbers they are
The strictest test I’ve run on the registry was pointing it at my own marketing copy. Two fabrication-rate statistics I’d been quoting failed verification. One couldn’t be traced to any published study; the other attributed a specific GPT-4 figure to a paper whose abstract doesn’t contain it. Both had to go, replaced with figures that survive verification. The pipeline threw out my own claims, which is exactly the point. The system doesn’t care whose plausible-sounding number it is.
None of this removes the human from the work. Someone still chooses the angle, structures the argument, writes and rewrites the prose, and decides which verified facts matter. What the pipeline removes is the ability to be confidently wrong about sources.
Build your own: the minimum viable version
You don’t need custom software to run an AI content pipeline this way. You need four rules. Research before writing: collect your statistics from live pages first, and record a verbatim quote and URL for each. Give the model an allow-list: paste your verified links into the prompt and forbid any source outside them. Audit mechanically: before publishing, check every link in the draft against your list, and anything else fails. And re-verify on refresh: when you update a piece, re-check that its sources still exist and still say what you claim.
What this looks like in practice
The registry is a spreadsheet. Four columns are enough: the claim, the source URL, the verbatim quote you copied from the live page, and the date you checked it. One row per statistic, filled in before you write a single sentence of the draft. A row you can’t complete is a claim you don’t get to make.
| Claim | Source URL | Verbatim quote | Checked |
|---|---|---|---|
| 18% of GPT-4’s citations in AI-generated papers were fabricated | nature.com/articles/s41598-023… | “just 18% of the GPT-4 citations are fabricated” | 2026-08-04 |
The allow-list is a paragraph in your prompt. Something like:
Write this section using only the sources listed below. Copy each link exactly as given, anchor text included. Do not cite, name, or link any source that is not on this list. If you need a fact you cannot support from these sources, make the argument without a statistic.
The audit is fifteen minutes with your browser’s find function. Search the draft for “http” and confirm every link appears in your sheet; anything that doesn’t gets verified on the spot or cut. Then search for “research shows”, “studies indicate” and “experts agree”: any hit without a link in the same sentence gets a source or gets rewritten as a direct argument.
And re-verification is a calendar entry. When a piece comes up for a refresh, open its sheet and click every URL. Dead pages and changed claims lose their row, and the copy updates with them.
Everything else in Optix is automation layered on those four rules. The rules are what make fabrication impossible; the software just makes them scale.