AI and Evidence Synthesis: Where It Helps, Where It Can't Replace Methodology
On this page
An AI evidence review tool now touches nearly every stage of evidence synthesis -- title/abstract screening assistants, extraction helpers, summarization tools that claim to turn a stack of PDFs into a literature review overnight. Some of this is genuinely useful. Some of it introduces exactly the kind of error a methods reviewer is trained to catch.
Example PRISMA flow diagram, showing how records narrow to included studies.
Here's an honest breakdown of where AI tools currently earn their place in a systematic review workflow, and where they don't.
Where AI genuinely helps
**Screening triage.** Machine-learning screening tools (like ASReview or the classifiers built into Covidence and Rayyan) can rank abstracts by likely relevance, letting reviewers work through the most probable includes first. This doesn't replace independent dual screening -- it just reorders the queue so obvious excludes get processed faster.
**Deduplication.** Removing duplicate records across multiple database exports is a mechanical task AI tools handle reliably, freeing up reviewer time for judgment calls.
**Search string drafting.** Large language models can draft a reasonable first pass at Boolean search terms for a topic, which a search strategist can then refine and test. Useful as a starting point, not as a finished search strategy.
**Reference formatting and citation checking.** Genuinely low-risk, high-time-savings territory.
Where AI tools fall short -- and why it matters
**Full-text screening and inclusion decisions.** Whether a study meets your PICO criteria often depends on nuance the abstract doesn't capture -- comparator details buried in the methods section, an outcome measured at a timepoint your protocol didn't specify. Current AI tools miss this kind of nuance regularly, and a missed inclusion or a wrongly included study changes your entire pooled estimate.
**Risk-of-bias assessment.** RoB 2 and ROBINS-I both require reading a study's actual methods critically -- was allocation concealment genuinely adequate, or does the paper just claim it was? This is a judgment call, not a pattern-match, and it's one of the areas where AI-generated risk-of-bias ratings have been shown to disagree substantially with expert reviewers.
**Data extraction from complex tables.** AI tools frequently misread multi-arm trial tables, subgroup breakdowns, or studies reporting the same outcome in different units. An extraction error here doesn't just affect one row -- it propagates into your meta-analysis.
**Full review generation.** Tools that claim to generate a complete systematic review from a topic prompt are, at best, producing a narrative summary dressed up as a systematic review. They skip protocol registration, they don't run a real reproducible search, and they can't perform genuine dual-reviewer screening. A journal's methods reviewer will catch this immediately -- and using one puts your academic integrity at risk in a way that's hard to walk back.
The honest framing
AI tools are good at compressing the mechanical parts of a systematic review -- the parts that were always tedious rather than intellectually demanding. They're not yet reliable for the parts that require judgment: deciding what counts as evidence, weighing risk of bias, and reconciling conflicting findings into a defensible synthesis.
That's also roughly where we draw our own line. We consult on methodology, we edit and advise, and we never author a review under our own name -- whether that review was drafted by a person or generated by a tool. The judgment calls stay with you and your named authorship, because that's what a journal's reviewers -- and your own academic integrity -- actually require.
If you're weighing which parts of your review workflow are safe to speed up with AI tools and which need a human second opinion, that's a conversation worth having before you start screening, not after.
**[Get a quote](/get-a-quote)** to talk through your workflow before you start screening.
This is exactly what we help with
Screening & Data Extraction →
[3] Screening & extraction
Get a Quote →