Systematic Reviews

Inter-Rater Reliability in Systematic Reviews: Understanding Kappa Statistics

July 27, 2026·Dr. Lauren Ito·5 min read
On this page

When two reviewers independently screen the same records, measuring how consistently they agree matters for demonstrating your screening process was genuinely reliable, not just nominally dual-conducted. Cohen's kappa is the standard statistic for this, and understanding what it actually measures, beyond a simple percentage agreement, changes how you should interpret and report it.

Records identifiedn = 2,314After deduplicationn = 1,802Title/abstract screenedn = 1,802Full-text assessedn = 96Studies includedn = 34Excluded (full-text)n = 62

Example PRISMA flow diagram, showing how records narrow to included studies.

Why raw percentage agreement is misleading

If two reviewers are screening records where 90 percent are obviously ineligible at a glance, they could agree 90 percent of the time simply by both correctly identifying the obvious exclusions, even with genuinely poor agreement on the harder, more ambiguous borderline cases that screening reliability actually needs to demonstrate. Raw percentage agreement doesn't account for this chance-driven baseline agreement, which is exactly the gap Cohen's kappa is designed to address.

What kappa actually calculates

Cohen's kappa compares your observed agreement between two reviewers against the level of agreement that would be expected purely by chance, given each reviewer's individual pattern of inclusion and exclusion decisions. A kappa of 0 indicates agreement no better than chance alone; a kappa of 1 indicates perfect agreement beyond what chance would produce.

Interpreting specific kappa values

Conventional interpretive bands, though somewhat arbitrary conventions rather than fixed scientific thresholds, generally treat kappa values below 0.20 as slight agreement, 0.21 to 0.40 as fair, 0.41 to 0.60 as moderate, 0.61 to 0.80 as substantial, and above 0.80 as almost perfect agreement. A systematic review reporting a kappa in the moderate range or below for its screening process should prompt genuine reflection on whether eligibility criteria were sufficiently clear, rather than being reported without comment as though any positive kappa value demonstrates adequate reliability.

When to actually calculate kappa

Kappa is most commonly calculated for title and abstract screening, often reported for the full screened set or for a defined proportion when full dual screening wasn't feasible throughout. It can similarly be calculated for full-text screening decisions and, less commonly but still usefully, for risk-of-bias judgments between two independent appraisers.

What a low kappa value actually tells you

A low kappa doesn't necessarily mean your eligibility criteria are wrong -- it might mean they were genuinely ambiguous as written, or that reviewers interpreted a specific criterion differently despite the wording seeming clear to whoever drafted it. Investigating which specific types of disagreement drove a low kappa, rather than simply reporting the number, often reveals a concrete, fixable issue in how eligibility criteria were communicated to the screening team.

Kappa's known limitations

Kappa can behave unexpectedly when the underlying prevalence of eligible versus ineligible records is very skewed in either direction, sometimes producing a low kappa value even when raw agreement is quite high, a phenomenon known as the kappa paradox. Being aware of this limitation, and considering supplementary agreement statistics or a qualitative review of actual disagreements in cases of very skewed eligibility rates, gives a more complete picture than relying on kappa alone in this specific situation.

Reporting kappa in your manuscript

State your kappa value explicitly in your methods section, ideally alongside the number of records it was calculated across, and briefly note whether this reflects full dual screening or a representative sample. If your kappa falls in a lower range, briefly addressing how disagreements were resolved and whether this affected your final included study list adds useful transparency beyond simply reporting the number itself.

Weighted kappa for non-binary judgments

For decisions with more than two possible categories, such as a three-point risk-of-bias judgment of low, some concerns, or high risk, weighted kappa accounts for the fact that a disagreement between adjacent categories is less serious than a disagreement between the two extreme categories, providing a more nuanced reliability measure than standard kappa, which treats all disagreements as equally serious regardless of category.

A practical takeaway

Calculating and reporting kappa, rather than a raw percentage agreement or simply asserting dual screening occurred without any reliability statistic at all, is expected practice for a methodologically rigorous systematic review, and a low kappa value is genuinely useful information worth investigating and reporting honestly rather than a result to downplay or omit from your final manuscript.

Improving kappa through better protocol clarity

Where piloting your screening criteria on a small sample reveals a lower-than-expected kappa, revising your eligibility criteria wording to address the specific points of disagreement, then re-testing on a fresh small sample, often meaningfully improves subsequent agreement, much like the piloting process described elsewhere for data extraction forms specifically. This kind of iterative refinement before full screening begins is considerably more efficient than discovering agreement problems only after screening thousands of records.

Kappa as an ongoing check, not just a one-time report

For very large reviews with an extended screening period, periodically recalculating kappa on a fresh sample partway through, not just once at the very end, can catch a gradual drift in how consistently reviewers are applying eligibility criteria over an extended screening timeline, before this drift meaningfully affects your final included study list. This kind of ongoing vigilance, though a small additional step, meaningfully strengthens a large review's claim to consistent, reliable screening across its entire, often lengthy, duration. This ongoing check costs relatively little time relative to a full screening pass, and it provides genuine, concrete reassurance that your dual-screening process remained reliable from its first record through its last. This kind of sustained, verifiable consistency is exactly the standard a rigorous systematic review's screening process is expected to meet, and demonstrating it explicitly strengthens a review's overall methodological credibility considerably. Teams that build this kind of periodic verification into their standard practice, rather than treating a single initial pilot as sufficient for an entire lengthy project, consistently produce more defensible, higher-quality final screening results. This modest, ongoing investment of time reflects the same broader principle that runs throughout rigorous systematic review methodology: verifying consistency actively, rather than simply assuming it holds throughout a long and demanding project, giving the whole team genuine confidence in the reliability of their final included study list.

#inter-rater reliability#Cohen’s kappa#screening