An AI-assisted draft can be fluent while combining a real peptide name with the wrong sequence, a plausible but nonexistent citation or an exaggerated interpretation. Fact-checking needs to inspect those individual claims rather than judge the text by how confidently it reads.
Check that each cited work exists and matches
Walters and Wilder’s 2023 study examined 636 references in 84 generated literature reviews using GPT-3.5 and GPT-4. It found fabricated citations and substantive errors in some references to real works. Those findings concern the systems and tasks studied, not an error rate for every current tool.Walters and Wilder (2023) — Fabrication and errors in the bibliographic citations generated by ChatGPT (opens in a new tab)
For each reference in a draft, locate the publisher or recognised bibliographic record and compare title, authors, year and identifier. A working link can still lead to a different paper from the one named in the text.
Then read the relevant source content. A real reference does not establish that the associated sentence is supported; the draft may attach the citation to a claim the paper never makes.
Record access limits. If only an abstract was inspected, do not describe the review as a full-text verification or invent details that would require access to methods, figures or supplementary files.
Check the entity before the conclusion
| Risk | Concrete check |
|---|---|
| Name conflation | Compare sequence, modification and molecular form where documented |
| Study conflation | Match model, endpoint and experimental material to the cited paper |
| Number drift | Compare value, unit, denominator and reporting basis |
| Evidence inflation | Separate hypothesis, observation and broader interpretation |
| Invented business fact | Check against an authorised business record |
A peptide family name, a fragment and a modified analogue may appear close together in training material or search results. The draft must not merge them simply because their names resemble one another.
Likewise, a supplier certificate and a mechanistic paper support different kinds of statements. Check that the draft does not treat a report’s analytical conformity as evidence for a biological effect.
If a sequence or identifier cannot be established, preserve the uncertainty. Replacing an empty field with a confident guess makes the document look more complete while reducing its reliability.
Recalculate and inspect the reasoning
Reperform arithmetic from the stated inputs, including unit conversions and percentage denominators. Check whether the calculation assumes a molecular form or quantity basis that the source does not establish.
For an original hypothetical example, a draft might claim that a change from 20 to 25 is a 5% increase. The absolute difference is 5 units, while the increase relative to 20 is 25%. Both the operation and its wording need correction.
Check qualitative claims with equal care. Terms such as selective, stable or validated need a defined comparison and supporting context; they should not be accepted because they sound like ordinary scientific vocabulary.
Keep a review record and resolve failures
Mark each material claim as supported, revised, unresolved or removed, with a source location where appropriate. This makes the remaining work visible and prevents a polished draft from being mistaken for a finished evidence review.
Check the corrected version again where edits affect connected statements. Narrowing a paragraph while leaving an overstated title or takeaway can preserve the original problem in the most prominent part of the article.
Describe AI assistance and editorial checking honestly under the applicable publication requirements. Do not claim expert review, independent laboratory verification or exhaustive literature coverage unless those processes actually occurred.
The goal is an accountable scientific explanation whose important statements survive inspection. AI can assist drafting and organisation, but responsibility for the final claims remains with the people choosing to publish or rely on them.
Sources and further detail
- Walters and Wilder (2023) — Fabrication and errors in the bibliographic citations generated by ChatGPT (opens in a new tab)
Original study abstract and methods inspected. Historical models and task scope retained; no error-rate estimate is transferred to current systems. The 20-to-25 calculation and checking matrix are original examples.
Sources checked 20 September 2026. Editorial examples are illustrative unless a supplied document or published study is explicitly identified. This article has not undergone independent scientific peer review.