Skip to main content

Screening papers for relevance

Four hundred titles is a manageable afternoon if you screen. It is an impossible week if you read.

How do you screen papers for relevance quickly?

Write the inclusion rule before you look at anything, as a short list of conditions a paper must meet. Then screen in two passes: title and abstract against the rule, then full text on whatever survived. Decide include, exclude or uncertain in seconds, record the reason for each exclusion, and never start reading during the first pass.

Write the rule before you look

The single most effective change is deciding what counts as relevant before seeing the results. A rule written afterwards drifts to fit whatever you found, and drift is invisible from the inside — by paper two hundred you are applying a different standard than at paper ten, and you will not notice.

A usable rule is four to eight conditions, each checkable from an abstract, each phrased so the answer is yes or no. Include the exclusions explicitly: not a conference abstract, not a case report, not before 2015, not an animal model.

Test the rule on ten papers you already know should be included, and ten you know should not. If it misclassifies any of them, fix the rule now. Fixing it at paper two hundred means rescreening everything.

Two passes, not one

PassInputDecisionTime per record
1Title, then abstract if the title is ambiguousInclude / Uncertain / Exclude5-20 seconds
2Full text of everything not excludedInclude / Exclude, with a recorded reason2-10 minutes

The first pass is deliberately generous: when in doubt, mark uncertain and move on. Excluding a relevant paper in pass one is unrecoverable, because you will never look at it again. Carrying a doubtful paper into pass two costs a few minutes.

The discipline that makes pass one fast is refusing to read. If you find yourself following an interesting argument in an abstract, you have stopped screening. Note it, mark it uncertain, move on.

Read abstracts in the right order

Abstracts are written front to back and are most efficiently screened out of order, because the fields that exclude a paper are rarely at the start.

  1. Population or subject — wrong population excludes immediately and is usually stated in one clause.
  2. Study design or method — the second most common exclusion, and also quick to check.
  3. Intervention or mechanism — what was actually done.
  4. Outcome — last, because it is the least reliably reported and the most likely to require the full text.

Roughly two thirds of exclusions are settled by the first two checks. Reading in that order rather than top to bottom is most of the speed difference between screening and reading.

Calibrate against a second reader

Dual independent screening is a requirement in systematic reviews and a good idea everywhere else. Two people screen the same set, then compare, and disagreement is resolved by discussion or a third reader.

Its value is not only catching errors. Disagreement in the first fifty records almost always reveals that the inclusion rule is ambiguous — and fixing the rule then is far cheaper than discovering the ambiguity at the end.

If you are working alone, approximate it: screen fifty records, do something else for a day, rescreen the same fifty blind, and compare. Your own disagreement rate is a usable measure of how well specified the rule is.

Record exclusion reasons

At full-text stage, every exclusion needs a recorded reason drawn from a fixed list — wrong population, wrong design, no relevant outcome, duplicate, no full text available. PRISMA requires this, and it is useful even when nothing requires it.

The reason list is also a diagnostic. If most exclusions are wrong population, the search blocks were too loose. If most are no relevant outcome, the question was underspecified. Reading the tally tells you which.

Where automation helps and where it does not

  • Deduplication across databases: fully automatable and should always be automated.
  • Priority ranking, so likely-relevant records surface first: reliable, and it makes the tail of a large screen much cheaper.
  • Assisted classification with a human confirming every exclusion: acceptable in many workflows, and increasingly common.
  • Automatic exclusion with no human review: not acceptable in anything claiming completeness, because a false exclusion is invisible and unrecoverable.

The safe rule is that a machine may reorder the queue and may propose, but a person decides every exclusion. Ranking changes how long screening takes; it must not change what ends up excluded.

Questions

How long should screening one abstract take?
Five to twenty seconds in the first pass. If it is consistently taking longer, you are reading rather than screening, or the inclusion rule has conditions that cannot be checked from an abstract and should be moved to the full-text pass.
What is a good inter-rater agreement for screening?
A Cohen kappa above about 0.6 is usually treated as acceptable and above 0.8 as strong, though the thresholds are conventions rather than rules. Low agreement early is normal and is a signal to clarify the inclusion rule before continuing, not to replace a reader.
Should I exclude papers I cannot access?
Not at screening. Record them as awaiting retrieval, then try Unpaywall, the repository version, interlibrary loan and the corresponding author. Papers excluded only because you could not obtain them must be reported as such, because that is a limitation rather than a decision.
Can I use AI to screen abstracts?
For ranking and for a first-pass suggestion, yes, and it saves substantial time on large sets. Not for unreviewed exclusion: a false exclusion never resurfaces, so a person should confirm every record that gets dropped.

Last updated 2026-08-25. Part of the litscout literature search guides.

We use only the cookies needed to run the service — signing you in and keeping the session secure. No analytics, no advertising, no third-party trackers. Privacy Policy