Fixed-Effect vs. Random-Effects Meta-Analysis: How to Choose
On this page
Choosing between a fixed-effect and a random-effects model is one of the most consequential decisions in a meta-analysis, and it is not a matter of preference. It rests on a specific assumption about your included studies, and that assumption is testable before you commit to either model.
What the fixed-effect model actually assumes
A fixed-effect model assumes every included study is estimating exactly the same true effect, and any variation you observe between study results is due entirely to sampling error -- random noise from having a finite sample size. Under this assumption, larger studies get proportionally more weight in the pooled estimate, because their sampling error is smaller and their estimate of the single true effect is more precise.
This assumption is rarely defensible in practice. It essentially requires that every study used an identical population, an identical intervention, an identical comparator, and an identical outcome definition. A fixed-effect model applied to a genuinely diverse set of studies produces a pooled estimate and confidence interval that are both too narrow and not meaningfully interpretable as "the" effect.
What the random-effects model assumes instead
A random-effects model assumes each study is estimating its own true effect, and these true effects are themselves drawn from a distribution of related but not identical effects. The model estimates both the average of that distribution and how much the true effects vary across studies (captured by tau-squared). This is a more realistic assumption for most systematic reviews, where studies differ in population characteristics, intervention delivery, follow-up duration, or outcome measurement -- even when they are answering a recognizably similar question.
Random-effects models generally produce wider confidence intervals than fixed-effect models on the same data, because they are honestly accounting for between-study variation rather than assuming it away. Smaller studies also carry relatively more weight under random-effects than they would under fixed-effect, since the model is not simply rewarding precision about a single shared truth.
The heterogeneity check that should inform your choice
Before committing to either model, run a heterogeneity assessment. The Cochran's Q test gives a formal significance test for whether observed variation exceeds what sampling error alone would explain, though it is known to have low power with a small number of studies. I-squared expresses the percentage of total variation attributable to genuine heterogeneity rather than chance, with rough conventional bands of low (around 25%), moderate (around 50%), and substantial (around 75%) heterogeneity -- though these thresholds are guidelines, not hard cutoffs, and should be interpreted alongside the actual clinical or methodological diversity of your included studies.
Meaningful heterogeneity, whether flagged by a significant Q test, a substantial I-squared, or simply visible diversity in study populations and designs, is a signal toward random-effects. Low heterogeneity combined with genuinely comparable studies can support a fixed-effect approach, though many methodologists now default to random-effects as the more conservative and generally applicable choice regardless of the heterogeneity statistics, precisely because the assumption of a single true effect is so rarely justified.
Why this decision draws reviewer scrutiny
A methods reviewer reading your statistical section will check whether your model choice is consistent with your reported heterogeneity statistics. Reporting substantial I-squared and then presenting a fixed-effect pooled estimate without justification is one of the most common statistical objections raised on submitted meta-analyses, because it suggests the model was chosen for convenience rather than because it matched the data.
The model choice should be specified in your protocol before you see the pooled results, for the same reason your eligibility criteria should be specified in advance -- to prevent the appearance that you selected the model that produced the most favorable or most precise-looking result.
What to do when heterogeneity is genuinely high
Very high heterogeneity, even under a random-effects model, can indicate that pooling is not the right choice at all. If I-squared is very high and driven by a small number of clear outlier studies, consider a sensitivity analysis excluding those studies to see how much the pooled estimate shifts. If the heterogeneity reflects genuinely different populations or intervention variants, a subgroup analysis or meta-regression exploring the source of that variation is often more informative than forcing a single pooled number that averages over meaningfully different effects.
In some cases, the honest conclusion is that the included studies are too heterogeneous to pool meaningfully, and a narrative synthesis presenting the range of findings is more defensible than a misleading single pooled estimate. This is a legitimate methodological outcome, not a failure of the review.
A practical decision framework
Specify your intended model in the protocol before running the analysis. Run the heterogeneity assessment -- Q test and I-squared at minimum -- as a standard step regardless of which model you expect to use. Choose random-effects as the default unless you have specific, pre-specified justification for fixed-effect, such as a genuinely narrow and homogeneous set of studies. Report the heterogeneity statistics alongside the pooled estimate either way, since a reviewer will expect to see them regardless of which model you ultimately report. If heterogeneity is very high, consider whether subgroup analysis, meta-regression, or narrative synthesis better serves the actual state of the evidence than a single pooled number.
A worked illustration
Consider five hypothetical trials of similar size measuring the same intervention, with individual risk ratios ranging from 0.65 to 0.95 -- a real spread, but all favoring the intervention. A fixed-effect model, assuming a single true effect, would produce a relatively narrow pooled confidence interval, reflecting only sampling error. A random-effects model applied to the same five trials, honestly incorporating the between-study spread as genuine variation rather than noise, would produce a visibly wider confidence interval around a similar central estimate. Neither number is "wrong" -- they answer subtly different questions, and only one of them is asking the question your heterogeneity statistics actually justify asking.
Getting this choice right, and documenting the reasoning behind it, is one of the clearest signals to a peer reviewer that your meta-analysis was conducted with genuine statistical care rather than run through a template.