Beyond Human-Likeness: Mapping the Scientific Critique Profiles of LLMs and Human Reviewers
Yunhan Yang, Mike Thelwall, Guoxiu He
Abstract
Large language models (LLMs) are increasingly discussed as tools for peer review, but their value is often assessed through human-likeness, perceived usefulness, or textual overlap with reviewer comments. This study shifts attention from whether LLMs resemble human reviewers to what functions of scientific critique they perform. Using ICLR 2025 peer-review data, we compare human reviews with LLM reviews generated under baseline and expert prompts. We operationalize scientific critique through two review acts, weakness critique and scientific questioning, and annotate point-level review text using five theory-guided frameworks: Anderson's knowledge types, Toulmin's argumentation model, Graesser's question depth, SOLO cognitive complexity, and Hattie's feedback functions. The results reveal a differentiated critique profile. Human reviews placed greater emphasis on scientific framing and revision guidance, more often identifying higher-order weaknesses and asking questions oriented toward improvement. LLM reviews showed higher rates of explanatory depth, integrative reasoning, and explicit argument structuring. Expert prompting did not make LLM critique uniformly more human-like; it partially narrowed some gaps but mainly amplified LLM-specific tendencies toward integration and formal argumentation. These findings show that LLM-assisted peer review changes the functional composition of review text, making it important to distinguish LLM-amplified critique from areas requiring human prioritization and accountable judgement.
Create a lesson
Related papers
Guiding LLM Peer Reviewers: The Impact of Score Anchors on Review Evidence and Accuracy
Judita Preiss, Yunhan Yang
Citing Less Critically: LLMs Reshape the Rhetoric and Reach of Scientific Citation
Yixuan Liu, Lin Chen, Zhuoqi Liu et al.
The zbMATH Open Knowledge Graph: Tracing Centuries of Mathematical Research
Yuni Susanti, Moritz Schubotz
Do Large Language Models Favour Any Research Topics?
Mike Thelwall
Tracing high-profile attention to questionable research as a case for funder due diligence
Federica Silvi, Leslie D. McIntosh
SoniMet - A tool for sonifying and visualizing the performance of single researchers
Tim Waterfield, Lutz Bornmann