Auto-judge: a cross-task benchmark for comparing LLM judges for citation-grounded RAG systems

Farzi, N. orcid.org/0009-0000-3297-8888, Hagen, T. orcid.org/0009-0000-4854-7249, Yang, E. orcid.org/0000-0002-0051-1535 et al. (7 more authors) (2026) Auto-judge: a cross-task benchmark for comparing LLM judges for citation-grounded RAG systems. In: SIGIR '26: Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval. 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2026), 20-24 Jul 2026, Melbourne, Australia. . ACM, pp. 3159-3166. ISBN: 9798400725999.

Abstract

Metadata

Item Type: Proceedings Paper
Authors/Creators:
Copyright, Publisher and Additional Information:

© 2026 Owner/Author. This work is licensed under a Creative Commons Attribution- 4.0 International License. https://creativecommons.org/licenses/by/4.0/

Keywords: llm-as-a-judge; evaluation; retrieval-augmented generation
Dates:
  • Published (online): 19 July 2026
  • Published: July 2026
Institution: The University of Sheffield
Academic Units: The University of Sheffield > Faculty of Engineering (Sheffield) > Department of Computer Science (Sheffield)
Date Deposited: 22 Jul 2026 10:46
Last Modified: 22 Jul 2026 10:56
Status: Published
Publisher: ACM
Refereed: Yes
Identification Number: 10.1145/3805712.3808601
Related URLs:
Open Archives Initiative ID (OAI ID):

Download

Export

Statistics