Skip to main content

What AI can and cannot do in a literature review

The useful question is not whether to use these tools. It is which specific steps they are trustworthy for.

Can AI write a literature review?

It can draft one, and the draft will be wrong in ways that are hard to see. AI is genuinely good at retrieval, ranking, deduplication and first-pass screening. It cannot assess methodological quality, cannot reliably detect retractions, and cannot support a completeness claim, because a similarity ranking is not a reproducible search.

Split the task before judging the tool

Literature review is not one task. It is at least seven, and current tools are excellent at some and unreliable at others. Most disappointment comes from applying a tool that is genuinely good at step two to step six.

StepAI reliabilityWhat to do
Find candidate workHighUse it. Semantic retrieval finds vocabulary you would not have searched.
Deduplicate across databasesHighAutomate fully.
Rank by likely relevanceHighUse it to order the screening queue.
First-pass screeningModerateUse as a suggestion. A person confirms every exclusion.
Extract stated data from a paperModerateUseful, and every extracted number needs checking against the source.
Assess methodological qualityLowDo it yourself. Tools do not read a methods section critically.
Synthesise and writeLow for claimsDraft structure, yes. Claims and citations, verify each one.

What it genuinely does well

  • Retrieval across vocabulary boundaries. Describing a problem in prose and getting back work that used entirely different terms is the clearest win, and it is exactly where Boolean search fails.
  • Ranking a large candidate set so the likely-relevant records surface first. This does not change what is in the set, only the order — which makes it safe as well as useful.
  • Deduplication, and reconciling metadata for the same work appearing in several databases.
  • Summarising what a paper says it did, as a screening aid, so you can decide whether to read it.
  • Surfacing structure — clustering a set of papers into the approaches they represent, which is a genuinely hard thing to do by hand across two hundred abstracts.

What it does badly, and why

The failures are not random. They follow from what these systems are.

  • Fabricated citations. A language model generating text will produce plausible references that do not exist. Tools grounded in a real index reduce this substantially but do not eliminate misattribution — a real paper cited for a claim it does not make is harder to catch than an invented one.
  • Methodological quality. Judging whether a study design supports its conclusion requires reading the methods against the claim. Summarisation reproduces what the abstract asserts, which is the thing you needed checked.
  • Retraction and correction status. Retracted papers continue to be cited for years, and few tools check retraction databases at retrieval time. Check the publisher page or Retraction Watch yourself for anything load-bearing.
  • Completeness. A top-k similarity ranking cannot state what it excluded, and nobody can re-execute it to verify. Any claim to have found the relevant evidence rests on a reproducible query, which is a Boolean one.
  • Recency at the edges. Index freshness varies, and the newest preprints are often the ones missing.

The dangerous failure is not the obvious one. An invented citation is caught the first time you click it. A real paper summarised slightly wrong, supporting a claim it does not quite make, survives into the draft and reads perfectly.

A workflow that respects the line

  1. Use semantic retrieval to explore and to harvest vocabulary. Read the results for terms, not yet for findings.
  2. Build and run the reproducible Boolean search from that vocabulary, in the databases whose coverage you need to claim. Log every string.
  3. Deduplicate automatically.
  4. Let a tool rank the screening queue. Screen it yourself, confirming every exclusion.
  5. Extract data with assistance, and verify every extracted number against the source document.
  6. Assess quality yourself. This step does not delegate.
  7. Write it yourself. Use a tool for structure and for finding the gaps in an argument, and verify every citation you keep — that it exists, that it says what you claim, and that it has not been retracted.

Disclosure

Most journals and universities now require disclosure of AI assistance, and the requirements differ. The consistent parts: AI cannot be an author, the human authors are responsible for everything in the text, and the use should be described specifically enough for a reader to judge it.

Used for screening support and for identifying candidate literature, with all included studies verified by the authors is a disclosure that tells a reader something. AI was used in preparing this manuscript is not.

Where litscout sits

litscout works on the retrieval side of that line. It models what a question is trying to establish, searches several paths including the citation graph, and explains why each result was returned — so the ranking can be argued with rather than taken on trust.

It does not assess methodological quality and does not check retraction status, and it says so in its own documentation. Its reports cite their sources and are a starting point for reading, not a substitute for it. For a review that has to defend completeness, it supplements a reproducible Boolean search rather than replacing one.

Questions

Will AI-written text be detected as plagiarism?
Detection tools are unreliable in both directions and should not be relied on by anyone. The real exposure is different: unverified AI text tends to contain citations that do not support the claim attached to them, and a reviewer who checks two references finds that immediately.
Can I use AI tools in a systematic review?
For screening support, deduplication, prioritisation and supplementary searching, yes, and several guidelines now address how to report it. Not as the primary search: the completeness claim depends on a query a reviewer can re-execute exactly.
Do AI research tools make up citations?
Ungrounded language models frequently do. Tools that retrieve from a real index return real papers, but can still attach a real citation to a claim the paper does not make. Verify that each source says what you are citing it for, not merely that it exists.
Which steps should never be delegated?
Assessing methodological quality, deciding what a paper actually establishes, confirming any exclusion, and verifying every citation you keep. These are the steps where an error is invisible in the finished text and changes the conclusion.

Last updated 2026-08-25. Part of the litscout literature search guides.

We use only the cookies needed to run the service — signing you in and keeping the session secure. No analytics, no advertising, no third-party trackers. Privacy Policy