5 Ways AI Deep Research Tools Invent Citations in Your Thesis
AI deep research tools produce references that look real and do not exist. Here are the five fabrication patterns that survive a quick read and surface in the viva.
The reference that looked perfect and did not exist
You ask a deep research tool to map the literature on your topic. Ten minutes later you have a tidy synthesis with thirty references, each formatted cleanly, each carrying a plausible author, journal, year, and page range. You paste a few into your bibliography. One of them does not exist. The authors are real, the journal is real, the article was never written.
This is the fabricated citation problem, and it has stopped being rare. A study published in the Lancet in 2026 tracked the rate of hallucinated references across more than two million papers and 97 million citations. The frequency rose from 1 in 2,828 papers in 2023 to 1 in 458 in 2025, and reached 1 in 277 during the first seven weeks of 2026. The tools that generate these references have improved at one specific thing: making the fakes look authentic.
For a thesis candidate the cost is not abstract. An examiner who finds one invented reference reads the rest of your bibliography differently. The question shifts from "is this argument sound" to "what else did this candidate not check." The five patterns below are the ones that pass a fast read.
Pattern 1: The plausible composite
The most dangerous fabrication is not invented from nothing. The tool stitches real elements into a citation that never appeared in print. A reviewer auditing NeurIPS 2025 submissions found a reference attributing the DeepFM model to Kingma and Ba at ICLR 2015. Kingma and Ba are real authors. ICLR 2015 is a real venue. What they actually published there was the Adam optimizer, not DeepFM, which came from a different team at IJCAI 2017. The citation is a composite of two real papers welded together.
A composite passes inspection because every individual part checks out. You recognise the authors. You recognise the venue. The year is reasonable. Only by retrieving the actual paper do you discover the combination is fictional. This is why scanning a reference list for "names I know" gives false confidence. The names being real is exactly what makes the fabrication convincing.
A psychology candidate citing a well-known author on attachment theory might accept a composite because the author genuinely writes about attachment. The specific paper, with its specific title and findings, was assembled by the tool. Thesisroom's citation guard checks each entry against Crossref, OpenAlex, and Semantic Scholar, so a composite fails verification even when its parts are individually real.
Pattern 2: The fabricated DOI
A DOI looks authoritative. It is a string of characters that resolves to a single registered object, so candidates treat its presence as proof the paper exists. Deep research tools generate DOIs in the correct format that resolve to nothing, or worse, resolve to a different paper entirely.
The format is the trap. A fabricated DOI follows the same pattern as a real one: a prefix, a slash, a suffix. Nothing about the string itself signals that it was invented. You can only tell by resolving it, and most candidates never do. They see the DOI, assume the reference is verified, and move on.
A reference with a round publication year (2020, 2015) and a clean DOI deserves extra scrutiny. Tools show a slight bias toward round years when fabricating, and a DOI that fails to resolve is a hard signal that the entry was generated rather than retrieved. Checking a DOI takes seconds. The failure to check is what reaches the examiner.
Pattern 3: The phantom from a sparse field
Fabrication is not evenly distributed. Tools invent more when their training data is thin for a subject. One study found GPT-4o fabricated references at 28 to 29 percent for less visible disorders such as body dysmorphic disorder, compared to 6 percent for major depressive disorder. Another team validating citations through CrossRef found hallucination rates above 80 percent for prompts referencing lower-income countries.
This pattern has a sharp implication for thesis candidates. If your research sits in a narrow subfield, a less-studied population, or a region outside the dominant literature, the deep research tool you used is more likely to have invented its sources, not less. The candidates most exposed to fabrication are often those doing the most original work, because originality means the literature is sparse.
A candidate studying a rare clinical condition, a minority language, or a specific regional policy cannot rely on the tool's output the way someone in a saturated field might. The thinner the genuine literature, the more aggressively the model fills the gap with plausible inventions. Every reference in a sparse field needs verification against the academic record before it enters your bibliography.
Pattern 4: The placeholder that slipped through
Some fabrications are not subtle at all. The NeurIPS audit found citations reading "Firstname Lastname," literal placeholder text that no human author would knowingly submit. Two such references passed peer review at one of the most respected conferences in the field. The lesson is not that reviewers are careless. The lesson is that nobody was checking citations systematically, so even trivially detectable errors survived.
A placeholder reaches your thesis when you trust the tool's output as a finished artifact rather than a draft. The deep research tool produced a synthesis, you accepted its structure, and a malformed entry rode along inside an otherwise clean list. The fix is mechanical: read every reference as a discrete object, not as part of a block you skim.
If a trivially broken citation can pass review at NeurIPS, an examiner reading your thesis closely will certainly find one in your bibliography. External examiners are expected to read the full thesis and submit a preliminary report before the viva. They have time to check what an automated pipeline missed.
Pattern 5: The real paper, wrong claim
The final pattern is the hardest to catch because the citation is real. The paper exists, the authors wrote it, the DOI resolves. The problem is that it does not say what the deep research tool claims it says. The tool retrieved a genuine paper and then attributed to it a finding the paper never reported.
This fabrication is invisible to any check that only verifies existence. Crossref will confirm the paper is real. The DOI will resolve. The only way to catch the mismatch is to read the source and confirm it supports the claim attached to it. In a viva, an examiner who knows the literature will recognise immediately when a cited paper is being made to say something it does not.
A candidate who writes "Smith (2019) found a significant effect" when Smith found no such effect has a real citation and a false claim. The reference passes automated verification and fails the moment an informed reader checks it. This is the gap between citation existence and citation accuracy, and it is where careful examiners concentrate.
What Thesisroom does with this
Thesisroom's citation guard extracts every entry in your bibliography and verifies it against Crossref, OpenAlex, Semantic Scholar, Retraction Watch, and DOAJ. It returns a traffic-light list: verified, flagged, or unverifiable. A composite fails because the specific combination does not match a real record. A fabricated DOI fails because it resolves to nothing or to a mismatched object. A phantom from a sparse field fails because no registry holds it.
The tool does not fabricate replacements or rewrite your references to hide gaps. Every flag points to a source you can check yourself. When a citation cannot be verified, Thesisroom's source finder helps you locate a verified alternative from the same registries, so you replace a phantom with a real paper rather than deleting the claim. The approach follows the same standard described in the integrity policy: every flag has a source.
Frequently asked questions
How common are fabricated citations from AI tools in 2026?
A 2026 Lancet study found the rate of hallucinated references rose from 1 in 2,828 papers in 2023 to 1 in 277 during early 2026. The tools have become better at making fabrications look authentic, with plausible author names, realistic dates, and convincing journal titles. The frequency is rising fastest in newer publications.
Can I trust a citation if the author and journal are real?
Not on its own. The most convincing fabrications are composites that combine real authors with a real journal and a paper that was never published. Verification means confirming that the specific entry, with its exact title and year, matches a real record, not just recognising the names involved.
Why do AI tools fabricate more in niche research areas?
Models fabricate more when their training data is sparse for a subject. Studies have documented fabrication rates near 28 percent for less-studied disorders and above 80 percent for prompts about lower-income countries. Candidates doing original work in narrow subfields face higher fabrication risk, not lower.
Does a working DOI mean the citation is verified?
A working DOI confirms the object exists but not that it supports your claim. Tools also generate correctly formatted DOIs that resolve to nothing or to a different paper. Checking that a DOI resolves is necessary but not sufficient; the source still has to say what you claim it says.
What happens if an examiner finds a fabricated reference in my viva?
One invented reference changes how the examiner reads the rest of your bibliography. The concern shifts from the quality of your argument to the reliability of your verification. Most outcomes still allow correction, but a fabricated citation discovered in the oral examination is an avoidable credibility cost.
Can a tool check whether a cited paper actually supports my claim?
Automated verification can confirm a paper exists and that its metadata matches. Confirming that the paper supports the specific claim attached to it still requires reading the source. Thesisroom flags entries that cannot be verified against the academic record, which is the first filter before you check the claim itself.
One thing to do before you submit
Treat every reference produced by a deep research tool as a draft entry, not a finished citation. Read your bibliography as a list of discrete objects and verify each one against the academic record before it stays in your thesis. The check that takes you a few hours is the check an examiner would otherwise do for you, in the room, out loud. Thesisroom's citation guard runs that verification across your full bibliography and returns a flag for every entry it cannot confirm.