Do Large Language Models Favour Any Research Topics?
Mike Thelwall
Abstract
Large Language Models (LLMs) can estimate the quality of published journal articles, potentially supporting human assessment when evaluations are needed. Whilst there are reasons to believe that LLMs may have biases in this role, there is no statistically strong evidence yet. The current article addresses this gap with an exploration of the types of articles that attract high or low LLM scores in 73,489 articles from 15 health and life sciences journals. Based on comparing the words in the titles and abstracts of higher and lower scoring articles for two LLMs in various ways, the results suggest that topics favoured by GPT-OSS-120B include viruses, genes and cells and its disfavoured topics include surveys, patients and students. It is not clear whether these patterns reflect underlying quality differences or AI biases, however. The same method found systematic differences between the topics favoured by GPT-OSS-120B and Gemma 3 27B, such as Gemma 3 27B giving relatively higher scores for machine learning research, proving that at least one of the two LLMs has AI bias. Finally, comparing the scores for full-text articles compared to scores for titles and abstracts also finds differences for both LLMs, showing that they both can exhibit AI bias for at least one of these two input types, and probably both. Overall, the results show that it is important to consider LLM biases when deciding whether to use them for research evaluation tasks.
Create a lesson
Related papers
Tracing high-profile attention to questionable research as a case for funder due diligence
Federica Silvi, Leslie D. McIntosh
SoniMet - A tool for sonifying and visualizing the performance of single researchers
Tim Waterfield, Lutz Bornmann
Measuring the Installed Base: Nordic Health Dataset Catalogues Against HealthDCAT-AP Release 7
Fabio Rovai
Gender and the Production of Research Impact
Sanger Wagner, Charles Rahal, Melinda C. Mills
A Comparative Evaluation of Digitization Pipelines for Historiographical Sources
Marina Gómez Rey, Patricia Callejo, Mario Muñoz-Organero et al.
COCI: Conference Organisers and Content Identifier
Angelo Salatino, Francesco Osborne, Alexis Vizcaino et al.