Does ChatGPT score research quality differently by gender?
Kayvan Kousha, Mike Thelwall
Abstract
Large Language Models (LLMs) are being considered for research evaluation, raising concerns about the introduction of AI bias. This study investigates whether ChatGPT research quality scores differ by first-author gender using 89,744 journal articles from the UK Research Excellence Framework (REF) 2021. Author information was withheld from ChatGPT to avoid direct gender bias. Nevertheless, male first-authored papers had slightly higher ChatGPT scores in most Units of Assessment (UoAs), especially in health, science and engineering-related subjects, and this pattern was often stronger for ChatGPT than for REF scores, based on a departmental-level proxy. Rank-based ChatGPT gains relative to REF scores were also more favourable for male first-authored papers in most UoAs, although the differences were generally small. Gender differences were not evident for solo research in the social sciences, arts and humanities, however. The male-favouring pattern for first-authored research was not explained by gender differences in writing styles, at least as reflected in abstract complexity. Some ChatGPT-REF differences may also reflect the departmental averaging process used to generate the REF proxy scores. Average ChatGPT scores may differ by first-author gender indirectly through other factors, such as field, topic, method, journal context or authorship structure. Thus, this is an additional reason to be cautious with AI-based research evaluation.
Create a lesson
Related papers
Shifting Research Funding Priorities under Geopolitical Pressure: Evidence from Estonia
Yunfeng Gao, Yang Ding
Geospatial Metadata Improves Discoverability by Connecting Datasets Across Scientific Disciplines
Daniel Ebanks, Devika Jain
Quantifying the impact of clinical-academic collaborations
Mohamad Zeina, Nick McNally, Karl S. Peggs et al.
Testing Our Foundations: Citation Trends, Errors, and Emerging Hallucinations in the Computing Education Literature
Paul Denny, Gweneth Barbre, Musa Blake et al.
Toward non-textual representation of social anthropology: Modeling cultures as knowledge graphs
Manolis Peponakis, Sarantos Kapidakis, Martin Doerr et al.
Geometric Signatures of Conceptual Reorganization: A Counterfactual Embedding Framework for Detecting Scientific Revolutions
Dimitris Ntounis, Ariel Schwartzman, Chris Chafe et al.