SUCRA Explained: Interpreting Network Meta-Analysis Rankings Correctly
On this page
- What SUCRA actually calculates
- Why SUCRA is appealing to report
- The core problem with interpreting SUCRA alone
- A concrete illustration of this problem
- What to check alongside every SUCRA ranking
- SUCRA and network geometry
- Presenting SUCRA responsibly in your manuscript
- An alternative or complementary approach: P-scores
- A practical rule for readers and authors alike
- SUCRA in the context of a growing evidence base
- Communicating SUCRA responsibly to clinical or policy audiences
SUCRA, the Surface Under the Cumulative Ranking curve, summarizes a network meta-analysis's ranking of multiple interventions into a single, easy-to-communicate number, which is exactly why it's become popular -- and exactly why it's also frequently over-interpreted without readers checking the actual uncertainty behind the ranking.
Example forest plot — each line is one study, the diamond is the pooled estimate.
What SUCRA actually calculates
For each intervention in your network, SUCRA calculates the probability that it ranks first, second, third, and so on among all compared interventions, then summarizes this entire ranking probability distribution into a single percentage between 0 and 100, where higher values indicate a greater likelihood of being among the better-performing interventions across the network.
Why SUCRA is appealing to report
A single number, easily displayed as a bar or ranked list, is far more immediately communicable to a general audience than a full table of pairwise comparisons with confidence intervals for every possible combination of interventions in your network. This communicative appeal is genuine and explains SUCRA's widespread adoption in published network meta-analyses.
The core problem with interpreting SUCRA alone
SUCRA reflects the probability of ranking well, not the magnitude or certainty of the actual difference between interventions. Two interventions can have meaningfully different SUCRA scores while their actual point estimates and confidence intervals overlap substantially, meaning the underlying evidence doesn't actually support confidently distinguishing between them, even though the SUCRA ranking visually suggests one clearly outperforms the other.
A concrete illustration of this problem
Imagine three interventions with SUCRA scores of 85, 70, and 40. This ranking looks decisive at a glance. But if the actual effect estimates underlying these rankings have wide, substantially overlapping confidence intervals, the true, clinically meaningful difference between the first and second-ranked interventions may not be statistically or clinically significant at all, despite the apparently clear SUCRA-based separation.
What to check alongside every SUCRA ranking
Always report the actual effect estimates and their confidence intervals for the specific pairwise comparisons your SUCRA ranking implies matter most, not the SUCRA score in isolation. A reader should be able to see both the ranking probability and the actual magnitude and certainty of the underlying differences driving that ranking, rather than being presented with only the more visually compelling summary number.
SUCRA and network geometry
A SUCRA ranking is only as reliable as the underlying network meta-analysis's transitivity assumption and the amount of direct versus indirect evidence supporting each comparison. An intervention ranking highly based largely on thin, indirect evidence through a sparse network deserves considerably more caution in interpretation than one ranking highly based on substantial, direct head-to-head evidence, even if their SUCRA scores appear numerically similar.
Presenting SUCRA responsibly in your manuscript
Report SUCRA rankings alongside, never instead of, a full league table showing pairwise effect estimates and confidence intervals for every relevant comparison in your network. A discussion section that references SUCRA rankings should explicitly address the underlying certainty and magnitude of the key comparisons driving those rankings, rather than treating the SUCRA percentage itself as the primary finding to communicate.
An alternative or complementary approach: P-scores
P-scores provide a broadly similar function to SUCRA using a frequentist rather than Bayesian analytical framework, and some methodologists consider them a useful complement or alternative depending on which underlying statistical approach your network meta-analysis uses. The same fundamental caution about not over-interpreting a single ranking summary number applies equally to P-scores.
A practical rule for readers and authors alike
Treat any SUCRA-based ranking claim with the same skepticism you'd apply to any other single summary statistic presented without its underlying uncertainty -- ask specifically what the actual pairwise estimates and confidence intervals show for the comparisons that matter most to your own decision, rather than accepting a ranked list at face value simply because it's presented with an air of statistical precision that a single percentage number can visually convey.
SUCRA in the context of a growing evidence base
As new trials are added to a network meta-analysis over time, SUCRA rankings can shift, sometimes substantially, particularly for interventions supported by relatively thin evidence within the network. Treating an early SUCRA ranking as a settled, final answer rather than a snapshot reflecting currently available evidence risks overstating how confidently a specific ranking should be trusted, especially for a rapidly evolving treatment area.
Communicating SUCRA responsibly to clinical or policy audiences
For readers using a network meta-analysis to inform a real treatment or policy decision, explicitly pairing any SUCRA-based ranking claim with a plain-language caveat about the underlying uncertainty, rather than presenting the ranking alone as a definitive answer, is a genuinely important communication responsibility for authors reporting this kind of analysis. Taking this responsibility seriously, rather than treating SUCRA as a self-explanatory number needing no further context, is part of what separates a genuinely useful network meta-analysis from one that inadvertently overstates its own certainty. Readers increasingly expect this level of interpretive care from authors reporting network meta-analysis findings, and providing it proactively, rather than waiting for a reviewer to request it, reflects genuinely careful, reader-focused scientific communication. This same principle of pairing a statistical finding with its appropriate caveats applies well beyond network meta-analysis specifically, reflecting a broader standard of responsible, reader-focused reporting throughout evidence synthesis work generally. Authors who consistently apply this standard across their published work tend to build a stronger, more trusted reputation among readers who rely on their statistical reporting to make genuinely informed decisions. This reputation, built consistently over time through careful, honest reporting practices, becomes a genuine professional asset that extends well beyond any single published analysis, extending its value well beyond the immediate results of any one specific review, and reinforcing this standard consistently is part of what keeps the whole field's reporting genuinely trustworthy over the long run, for every researcher who later builds on it in their own work, long after the original network meta-analysis was first published, as later researchers continue to build on and reference the same underlying comparisons, refining and extending the picture your original analysis first sketched out.