Thelwall, M. orcid.org/0000-0001-6065-205X (2026) In which fields do ChatGPT scores align more closely with research quality than do citation rates? Journal of Data and Information Science. ISSN: 2543-683X
Abstract
Purpose
Although citation-based indicators are widely used, they are not useful for recently published research, directly reflect only one of the three common dimensions of research quality, and have little value in some social sciences, arts and humanities. Large Language Models (LLMs) may address some of these weaknesses.
Design/methodology/approach
This article reports a science-wide assessment of the research quality scoring capability of ChatGPT-4o mini, ChatGPT-4o, and ChatGPT-5 mini. It correlates ChatGPT scores, averaged over 5 repetitions, with departmental average quality scores for 107,212 UK-based journal articles.
Findings
ChatGPT-4o is marginally better than ChatGPT-4o mini in most of the 34 field-based Units of Assessment (UoAs) tested. ChatGPT-4o scores have a positive correlation with research quality in 33 of the 34 UoAs, with the results being statistically significant in 31. ChatGPT-4o scores had a higher correlation with research quality than long term citation rates in 21 out of 34 UoAs and a higher correlation than short term citation rates in 26 out of 34 UoAs. The most substantial exception is Physics, for which citations are more useful. ChatGPT-5 mini has even stronger correlations overall and for departmental averages, it correlates more strongly with quality scores than do citations in 31 out of 34 UoAs, with correlations reaching 0.905.
Research limitations
All articles assessed are from the UK. The practical value of LLM scores for decision making is not assessed. Only the Normalised Log-transformed Citation Score (NLCS) was tested against ChatGPT rather than other citation rate indicators.
Practical implications
ChatGPT can be considered as a research quality indicator to support expert judgement in almost all academic fields.
Originality/value
The results give science-wide evidence that ChatGPT-4o mini, ChatGPT-4o, and ChatGPT-5 mini are competitive with citations as new research quality indicator sources, and technically better in most fields.
Metadata
| Item Type: | Article |
|---|---|
| Authors/Creators: |
|
| Copyright, Publisher and Additional Information: | © 2026 the author(s), published by De Gruyter on behalf of the Chinese Academy of Sciences This work is licensed under the Creative Commons Attribution 4.0 International License. (https://creativecommons.org/licenses/by/4.0/) |
| Keywords: | ChatGPT; large language models; research evaluation; scientometrics |
| Dates: |
|
| Institution: | The University of Sheffield |
| Academic Units: | The University of Sheffield > Faculty of Social Sciences (Sheffield) > School of Information, Journalism and Communication |
| Funding Information: | Funder Grant number UK RESEARCH AND INNOVATION UKRI1079 |
| Date Deposited: | 03 Aug 2026 07:59 |
| Last Modified: | 03 Aug 2026 07:59 |
| Status: | Published online |
| Publisher: | Walter de Gruyter GmbH |
| Refereed: | Yes |
| Identification Number: | 10.1515/jdis-2026-0058 |
| Open Archives Initiative ID (OAI ID): | oai:eprints.whiterose.ac.uk:244074 |
Download
Filename: 10.1515_jdis-2026-0058.pdf
Licence: CC-BY 4.0

CORE (COnnecting REpositories)
CORE (COnnecting REpositories)