100+methodology articlesWritersScribe
Systematic Reviews

A Systematic Review on Large Language Models in Research

July 23, 2026·James Okafor·5 min read
On this page

Large language models are increasingly applied within the research process itself -- assisting with literature screening, automating aspects of data extraction, supporting manuscript drafting, and even contributing to peer review -- and a systematic review examining this topic needs to commit to one specific application within the research workflow rather than attempting an overly broad survey of AI's role in research generally.

Narrowing to a specific stage of the research process

A review examining large language models' accuracy assisting with systematic review screening is asking a genuinely different question than one examining their use in manuscript drafting or peer review support, and each application raises distinct evidence requirements, distinct appropriate outcome measures, and distinct quality concerns worth addressing in a dedicated, tightly scoped review.

Screening assistance as a diagnostic-accuracy-structured question

Where your review addresses large language models' accuracy assisting with title and abstract screening specifically -- a topic of genuine, active interest given systematic review methodology's own reliance on this labor-intensive step -- this structures as a diagnostic accuracy question, comparing the model's inclusion and exclusion decisions against a human reviewer's decisions as the reference standard, following the same QUADAS-2-adjacent logic relevant to other AI diagnostic accuracy questions covered in this series.

Data extraction automation as a distinct evidence type

Research examining large language models' accuracy extracting structured data from primary studies represents a related but distinct question, typically measured through agreement rates between model-extracted and human-extracted data across a defined set of fields, requiring an extraction approach in your own review that captures this specific accuracy or agreement metric consistently across included studies.

Manuscript and writing support as a more qualitative question

Where your review addresses large language models' role supporting manuscript drafting or academic writing more broadly, evidence here is often more qualitative or survey-based -- researcher experiences, perceived quality changes, workflow efficiency -- than the more quantifiable accuracy questions relevant to screening or extraction assistance, and your review's methodology and appraisal approach should reflect this genuinely different evidence type.

This topic's unusually fast-moving evidence base

Given how recently and rapidly large language models capable of these research-assistance tasks have emerged, this specific systematic review topic is among the fastest-evolving covered in this series, and a living systematic review approach, or at minimum an explicit, prominent statement of your search date's currency limitations, is particularly warranted here.

Search strategy across research methodology and technical literature

This topic's relevant literature spans research methodology journals, informatics and text-mining venues, and increasingly dedicated venues discussing AI's role in scholarly publishing specifically, meaning a comprehensive search needs to look beyond conventional subject-matter databases toward this more specific, emerging intersection of literatures.

A genuinely reflexive consideration worth naming

Reviews on this specific topic occupy an interesting position -- a systematic review examining AI tools' role in conducting systematic reviews is, in principle, itself a candidate for applying the very tools it studies, and being transparent in your own methods section about whether and how you used any AI assistance in conducting your own review adds a genuine layer of relevant methodological transparency this particular topic specifically invites.

A practical starting point

Before beginning your search, identify precisely which stage of the research workflow your review addresses, confirm whether this structures as a diagnostic accuracy question, an extraction agreement question, or a more qualitative experience-focused question, and plan for this literature's genuinely rapid pace of change when setting your search date and considering whether an update or living review approach is warranted.

Considering the reviewer's own role transparently

Because this specific topic concerns tools that could, in principle, assist with conducting the very review examining them, being unusually explicit and transparent about your own review team's methodology, including any AI tool use in your own process, models the kind of transparency this literature itself is actively grappling with as a field.

A note on validation against genuinely independent human performance

Where your included studies compare large language model performance against human reviewer performance, confirming that the human comparison genuinely reflects independent, blinded performance, rather than a comparison that may have been influenced by prior exposure to the AI tool's own outputs, is a specific quality consideration worth checking carefully given this particular research question's somewhat unusual, self-referential structure.

A closing consideration on this topic's evolving relevance

As these tools become more deeply embedded in how research itself gets conducted, a well-scoped, honestly reported systematic review on this specific topic offers genuinely practical value to the broader research community navigating this transition in real time. That evolving relevance is part of what makes this a genuinely rewarding, if demanding, area to contribute careful systematic review work to right now.

A final word on modeling the transparency this field needs

As a reviewer working on this specific topic, your own transparent documentation of methodology and any AI tool use in your process becomes a small but genuine contribution to the broader normative conversation this field is actively having about responsible AI use in research, beyond just the substantive findings your review reports.

Considering the researchers who will build on your work

Other researchers navigating this same fast-moving topic will likely rely on your review as a starting point for their own work, and documenting your search strategy and decision-making with particular care, given how quickly this field's conventions and terminology continue to shift, makes your review a genuinely more durable and reusable resource for those who follow. That durability is a genuinely worthwhile goal in a field changing quickly enough that today's careful documentation may be tomorrow's most valuable reference point, worth building with the same care you would want a future researcher to extend to your own work. This extra care is warranted precisely because so many others will likely build directly on it in a field changing quickly enough that today's documentation can become tomorrow's essential starting point. Treating documentation this seriously is a small discipline with a genuinely lasting payoff.

#large language models#research methodology#systematic reviews