Fetching the paper…
Reading the bibliography…
The use of large language models (LLMs) for relevance assessment in information retrieval has gained significant attention, with recent studies suggesting that LLM-based judgments provide comparable evaluations to human judgments.
Reciprocal rank fusion outperforms condorcet and individual rank learning methods
Gordon V. Cormack, Charles L A Clarke, and Stefan Buettcher · 2009
Earlier work this paper cites.
Assessing top- k k preferences
Charles L. A. Clarke, Alexandra Vtyurina, and Mark D. Smucker · 2021
Earlier work this paper cites.
Perspectives on large language models for relevance judgment
Guglielmo Faggioli, Laura Dietz, Charles L. A. Clarke, Gianluca Demartini, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Benno Stein, and Henning Wachsmuth · 2023
Earlier work this paper cites.
Llms as narcissistic evaluators: When ego inflates evaluation scores
Yiqi Liu, Nafise Sadat Moosavi, and Chenghua Lin · 2023
Cited alongside, same era.
Large language models are not fair evaluators
Peiyi Wang, Lei Li, Liang Chen, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui · 2023
Cited alongside, same era.
LLMs can be fooled into labelling a document as relevant: best café near me; this paper is perfectly relevant
Marwah Alaofi, Paul Thomas, Falk Scholer, and Mark Sanderson · 2024
Cited alongside, same era.
A large-scale study of relevance assessments with large language models: An initial look, 2024a
Shivani Upadhyay, Ronak Pradeep, Nandan Thakur, Daniel Campos, Nick Craswell, Ian Soboroff, Hoa Trang Dang, and Jimmy Lin
Cited in the paper.
UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor, 2024b
Shivani Upadhyay, Ronak Pradeep, Nandan Thakur, Nick Craswell, and Jimmy Lin
Cited in the paper.
A comparison of methods for evaluating generative ir, 2024
Negar Arabzadeh and Charles L. A. Clarke · 2024
Closest in time.
Llm evaluators recognize and favor their own generations
Arjun Panickssery, Samuel R Bowman, and Shi Feng · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…