100+methodology articlesWritersScribe
Systematic Reviews

How to Conduct a Systematic Review on AI in Healthcare

July 1, 2026·Dr. Grace Mwangi·5 min read
On this page

Few topics are generating as many systematic review submissions right now as artificial intelligence applications in healthcare, and for good reason -- the underlying primary research is expanding faster than most reviewers can comfortably track. Conducting a rigorous systematic review on AI in healthcare specifically requires adapting standard methodology to a field that does not always fit cleanly into conventional intervention-review conventions.

Framing a genuinely answerable question

"AI in healthcare" as a topic is far too broad to review systematically on its own -- it spans diagnostic imaging algorithms, clinical decision support tools, administrative automation, patient-facing chatbots, and predictive risk models, each raising genuinely different methodological questions. A defensible review narrows this to a specific AI application type, a specific clinical context, and a specific outcome, framed in PICO terms just as any other systematic review would require: which specific AI tool or model class, applied to which patient population or clinical task, compared against which standard of care or human-performed baseline, measured against which specific outcome.

Why standard risk-of-bias tools often do not fit cleanly

Many AI-in-healthcare primary studies are neither randomized trials nor conventional observational studies in the sense RoB 2 or ROBINS-I were built around -- they are frequently algorithm validation studies, reporting a model's diagnostic accuracy against a reference standard rather than a comparative clinical outcome. For this kind of study, QUADAS-2, the tool built for diagnostic test accuracy research, is often the more appropriate fit than RoB 2, and reviewers new to this space commonly default to the wrong tool simply out of unfamiliarity with how AI validation studies are typically structured.

The PROBAST tool for prediction models specifically

Where your included studies involve AI models built to predict a clinical outcome -- readmission risk, disease progression, treatment response -- PROBAST, the Prediction model Risk Of Bias ASsessment Tool, is purpose-built for exactly this kind of study and addresses concerns specific to model development and validation that neither RoB 2 nor QUADAS-2 was designed to catch, including how the model was trained, whether validation used genuinely independent data, and whether performance metrics were reported completely.

Search strategy challenges specific to this field

AI-in-healthcare research is published across an unusually wide range of venues -- traditional clinical journals, computer science and engineering conferences, and preprint servers specifically serving the machine learning research community. A search strategy limited to conventional health databases like MEDLINE and Embase alone risks missing a meaningful share of the relevant literature, particularly technical validation studies published primarily in computer science venues. Searching arXiv and relevant machine learning conference proceedings, alongside standard health databases, is increasingly necessary for genuine comprehensiveness in this specific topic area.

Terminology inconsistency as a genuine search challenge

This field's terminology is still actively evolving, and the same underlying technique gets described inconsistently across the literature -- "machine learning," "deep learning," "artificial intelligence," and specific architecture names like "convolutional neural network" or "large language model" are sometimes used almost interchangeably by authors and sometimes distinguished carefully, depending on the publishing venue and the authors' own field. A comprehensive search strategy needs to include this full range of terminology rather than assuming consistent, standardized language across the literature.

Currency of your evidence base

Given how quickly the underlying AI models themselves evolve, a systematic review on this topic can become outdated unusually fast compared to a review on a more slowly evolving clinical intervention. Explicitly stating your search date prominently, and considering whether a living systematic review approach is warranted given this field's genuine pace of change, is worth deliberate consideration at the protocol stage rather than an afterthought.

Reporting model performance metrics consistently

AI validation studies report performance using a range of metrics -- sensitivity, specificity, area under the receiver operating characteristic curve, F1 score -- and your extraction form needs to capture whichever specific metrics your included studies report, with enough consistency across studies to support meaningful synthesis, rather than assuming every study reports the same standardized set of outcomes.

Getting methodology support for this specific topic area

Because AI-in-healthcare systematic reviews sit at the intersection of clinical methodology and technical AI literature that many reviewers are not equally fluent in both halves of, this is a topic area where methodology consulting genuinely pays for itself -- confirming your tool choice between RoB 2, QUADAS-2, and PROBAST early, and building a search strategy that genuinely covers both the clinical and technical literature, prevents a considerable amount of rework discovered only after screening is already underway.

A note on interdisciplinary team composition

Given how much this specific review type benefits from genuine fluency in both clinical methodology and technical AI literature, teams composed entirely of clinical researchers, or entirely of technical AI researchers, often benefit from deliberately pairing with someone from the complementary discipline, rather than attempting the full review from a single disciplinary background alone. This pairing is worth arranging explicitly at the protocol stage, not discovered as a gap only once appraisal or synthesis proves genuinely difficult.

Reporting AI-specific details in your final manuscript

Beyond standard PRISMA reporting, a systematic review on this topic benefits from explicitly reporting each included study's specific AI model type, training data characteristics where available, and validation approach, since these details matter considerably for readers assessing how applicable a given study's findings are to their own clinical context or research question. Treating these as standard, expected extraction fields, rather than optional extras, reflects genuine methodological maturity specific to this review type.

A closing consideration on this topic's practical value

A well-conducted systematic review here does more than summarize existing evidence -- it gives clinicians, health system leaders, and researchers a genuinely trustworthy reference point in a field where hype and rigorous evidence are not always easy to distinguish from the outside, making methodological care in this specific area particularly consequential. Approached with this level of care, a systematic review on AI in healthcare becomes a genuine resource, not just another entry in an already crowded literature.

#AI in healthcare#systematic reviews#methodology