+Low risk!Some concernsHigh riskQUADAS-2
Quality Appraisal

Newcastle-Ottawa Scale for Cohort and Case-Control Studies

August 15, 2026·Dr. Amara Chen·3 min read
On this page

Cochrane's RoB 2 covers randomized trials. ROBINS-I covers non-randomized intervention studies. But a large share of the observational evidence in fields like epidemiology, nursing, and public health comes from cohort and case-control studies, and the Newcastle-Ottawa Scale is the tool most reviewers reach for to appraise them.

What the Newcastle-Ottawa Scale actually assesses

The Newcastle-Ottawa Scale (NOS) uses a star-rating system across three domains, with slightly different criteria depending on whether you're appraising a cohort study or a case-control study.

**Selection** (up to 4 stars for cohort studies, 4 for case-control): checks how representative the exposed cohort or cases are, how the non-exposed cohort or controls were selected, and how exposure or case status was ascertained. A study using a community-representative cohort with a validated exposure measurement scores well here; one drawing from a narrow, unrepresentative sample does not.

**Comparability** (up to 2 stars): checks whether the study controlled for the most important confounding factors, and whether it controlled for additional relevant factors beyond that. This is the domain reviewers most often disagree on, since "most important confounder" is field-specific and requires genuine subject-matter judgment, not just a mechanical checklist pass.

**Outcome** (cohort studies) or **Exposure** (case-control studies), up to 3 stars: for cohort studies, this checks how the outcome was assessed, whether follow-up was long enough for outcomes to occur, and whether follow-up was adequate (typically requiring a stated threshold, such as 80 percent retention). For case-control studies, it checks how exposure was ascertained and whether the same ascertainment method was applied to cases and controls.

How scoring works, and its real limitation

A study can earn up to 9 stars total. Many reviews then apply a threshold, commonly treating 7 to 9 stars as low risk of bias, 4 to 6 as moderate, and below 4 as high, though there is no single universally agreed cutoff, and using one without justifying it is a common point of pushback from methods reviewers.

The scale's real limitation is that it produces a single summary score, which can obscure exactly where a study's weaknesses lie. A study can lose stars in comparability due to unmeasured confounding, or in outcome ascertainment due to a validated instrument, and both would show up as a similar overall reduction despite very different implications for how much you should trust the study's findings. For that reason, many methodologists recommend reporting star allocations by domain, not just the total, so readers can see exactly where the risk concentrates.

Common mistakes when applying NOS

**Not adapting the comparability criteria to the specific research question.** The scale gives no fixed list of which confounders qualify as "most important." Reviewers need to define this in advance, ideally in the protocol, based on subject-matter knowledge of the exposure-outcome relationship being studied.

**Applying it to randomized trials.** NOS is designed specifically for cohort and case-control studies, not trials. Using it on randomized evidence is a methodological mismatch that a reviewer will flag immediately.

**Treating a single reviewer's rating as final.** Like other risk-of-bias tools, NOS should be applied by at least two independent reviewers, with disagreements resolved by discussion, particularly given how much judgment the comparability domain requires.

**Skipping the accompanying guidance document.** The Newcastle-Ottawa Scale has separate, more detailed coding manuals for cohort and case-control studies that clarify what counts as an adequate answer for each item. Scoring from memory of the checklist alone, without the coding manual, is a common source of inconsistency between reviewers.

Where NOS fits into your overall risk-of-bias strategy

If your review includes a mix of study designs, you may need more than one tool: RoB 2 for any included trials, NOS for cohort and case-control studies, and potentially another tool for cross-sectional designs. GRADE certainty ratings should then draw on all of these consistently, downgrading for risk of bias based on the actual proportion of evidence coming from lower-scoring studies across whichever tools apply.

**[Get a quote](/systematic-reviews#risk-of-bias-assessment)** for a second opinion on your risk-of-bias appraisal, whichever tools your review requires.

#newcastle-ottawa scale