Quality Appraisal Tools Beyond Risk of Bias: AMSTAR, CASP, and More
On this page
- AMSTAR-2: appraising systematic reviews themselves
- CASP checklists: broad, practice-oriented appraisal
- QUADAS-2: diagnostic test accuracy studies
- NOS: the Newcastle-Ottawa Scale
- GRADE-CERQual: certainty for qualitative findings
- Choosing the right tool for your situation
- Applying more than one tool in a single review
- A quick way to check you've chosen correctly
- Appraisal tools continue to evolve
- Reporting appraisal results transparently
- Appraisal tools and non-English or grey literature sources
Risk of bias assessment, using RoB 2 or ROBINS-I, is the appraisal step most systematic reviewers are familiar with, but it is not the only quality appraisal tool relevant to evidence synthesis work. Several other instruments exist for genuinely different appraisal purposes, and knowing which applies to which situation matters for choosing correctly.
AMSTAR-2: appraising systematic reviews themselves
AMSTAR-2, A MeaSurement Tool to Assess systematic Reviews, is used specifically when your unit of appraisal is a systematic review itself rather than a primary study -- most commonly in umbrella reviews, where you're evaluating the methodological quality of the systematic reviews you're including. AMSTAR-2 assesses domains including protocol registration, search comprehensiveness, risk-of-bias assessment of the primary studies within each included review, and appropriateness of the review's own synthesis methods, producing an overall confidence rating from high to critically low.
CASP checklists: broad, practice-oriented appraisal
The Critical Appraisal Skills Programme provides a family of checklists covering various study designs -- RCTs, cohort studies, case-control studies, qualitative research, systematic reviews, and more -- widely used particularly in clinical and healthcare education settings. CASP checklists tend to be shorter and more accessible than some alternatives, making them a common choice for teaching critical appraisal skills, though this accessibility comes with somewhat less granular domain-by-domain structure than RoB 2 or ROBINS-I.
QUADAS-2: diagnostic test accuracy studies
For systematic reviews evaluating a diagnostic test's accuracy rather than an intervention's effectiveness, QUADAS-2 is the standard risk-of-bias tool, assessing domains specific to diagnostic accuracy research -- patient selection, the index test, the reference standard, and flow and timing through the study -- concerns that don't map onto RoB 2's intervention-trial-focused domains.
NOS: the Newcastle-Ottawa Scale
An older but still used tool for appraising non-randomized studies, particularly cohort and case-control designs, the Newcastle-Ottawa Scale uses a star-rating system across selection, comparability, and outcome or exposure domains. It has been substantially superseded by ROBINS-I in Cochrane-methodology reviews but still appears in some published literature and journal expectations, particularly outside formal Cochrane review production.
GRADE-CERQual: certainty for qualitative findings
Serving a function for qualitative evidence synthesis analogous to what GRADE does for quantitative systematic reviews, CERQual assesses confidence in individual qualitative review findings across methodological limitations, relevance, coherence, and data adequacy, producing a certainty rating that helps readers calibrate confidence in your qualitative synthesis conclusions.
Choosing the right tool for your situation
The choice is determined by what you're actually appraising and your review's methodology, not by preference among broadly similar-sounding options. A standard systematic review of RCTs uses RoB 2. The same review including observational studies also needs ROBINS-I for those studies specifically. An umbrella review appraises its included systematic reviews with AMSTAR-2. A diagnostic accuracy review uses QUADAS-2. A qualitative evidence synthesis uses CASP or an equivalent qualitative appraisal tool for individual studies, and CERQual for overall finding-level certainty.
Applying more than one tool in a single review
Reviews combining multiple evidence types or review levels sometimes genuinely need more than one appraisal tool applied to different components -- an umbrella review appraising both its included systematic reviews with AMSTAR-2 and, where relevant, drilling into primary study quality within a particularly important included review. This should be planned explicitly in your protocol and reported clearly in your methods section, specifying exactly which tool was applied to which category of included source.
A quick way to check you've chosen correctly
Before beginning appraisal, confirm your chosen tool was actually designed for your specific unit of appraisal -- a primary study of a specific design, a systematic review, a qualitative study -- and not simply the most familiar or most commonly cited tool in your general field. This small check prevents one of the more avoidable methodological mismatches in the entire systematic review process.
Appraisal tools continue to evolve
Newer or updated versions of these tools appear periodically -- AMSTAR-2 itself superseded the original AMSTAR, and further refinements continue to be proposed in the methodological literature. Checking that you're using the current, widely accepted version of whichever tool applies to your review, rather than an older version inherited from a template or previous project, is worth the few minutes it takes to verify against the tool developer's own current publication.
Reporting appraisal results transparently
Whatever tool or tools you use, presenting the actual item-by-item appraisal results, not just a single summary judgment, in a supplementary table gives readers the ability to see exactly which specific domains drove your overall quality assessment for each included source. This level of transparency is increasingly expected practice and makes your appraisal process auditable by anyone wanting to check your work against the same criteria you applied.
Appraisal tools and non-English or grey literature sources
Some appraisal tools assume a level of methodological reporting detail that grey literature sources or non-English publications translated for review don't always provide, which can make formal appraisal genuinely harder to apply consistently to these sources. Deciding in advance how you'll handle appraisal for sources with less complete reporting -- applying the tool as best you can with transparent notes on missing information, versus a modified approach -- is worth specifying in your protocol rather than deciding ad hoc when you encounter the first such source. Deciding this in advance also keeps your appraisal process consistent across your full included study list, rather than having your approach shift partway through simply because a difficult source happened to come up. Teams that document this decision clearly in their protocol also find it considerably easier to respond to a reviewer question about a specific hard-to-appraise source later, since the reasoning is already recorded rather than needing to be reconstructed from memory well after the appraisal work was completed. This kind of forward planning around difficult sources is a small, low-cost habit that consistently pays off later, particularly for reviews that include grey literature or non-English sources as a deliberate, disclosed part of their search strategy. Planning for this in advance, rather than reacting to it mid-review, keeps your appraisal process consistent and genuinely defensible throughout.