Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are increasingly deployed in both academic and industry settings to automate the evaluation of information seeking systems, particularly by generating graded relevance judgments.
A nugget-based test collection construction paradigm. In
Shahzad Rajput, Virgil Pavlu, Peter B Golbus, and Javed A Aslam. 2011 · 1948
Earlier work this paper cites.
Are Large Language Models Good at Utility Judgments?. In
Hengran Zhang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke, Yixing Fan, and Xueqi Cheng. 2024 · 1951
Earlier work this paper cites.
The significance of the Cranfield tests on index languages. In
Cyril W Cleverdon. 1991 · 1991
Earlier work this paper cites.
The first text retrieval conference (TREC-1) . Vol. 500
Donna K Harman. 1993 · 1993
Earlier work this paper cites.
Evaluating natural language processing systems: An analysis and review
Karen Sparck Jones and Julia R Galliers. 1995 · 1995
Earlier work this paper cites.
Efficient construction of large test collections. In
Gordon V. Cormack, Christopher R. Palmer, and Charles L. A. Clarke. 1998 · 1998
Earlier work this paper cites.
Variations in relevance judgments and the measurement of retrieval effectiveness. In
Ellen M. Voorhees. 1998 · 1998
Earlier work this paper cites.
Overview of the trec-8 web track. In
David Hawking, Ellen Voorhees, Nick Craswell, Peter Bailey, et al · 1999
Earlier work this paper cites.
Overview of IR tasks at the first NTCIR workshop. In
Noriko Kando, Kazuko Kuriyama, Toshihiko Nozue, Koji Eguchi, Hiroyuki Kato, and Souichiro Hidaka. 1999 · 1999
Earlier work this paper cites.
The PageRank citation ranking: Bringing order to the web
Lawrence Page. 1999 · 1999
Earlier work this paper cites.
Report on trec-9. In
Ellen M Voorhees. 2000 · 2000
Earlier work this paper cites.
Ranking Retrieval Systems without Relevance Judgments. In
Ian Soboroff, Charles Nicholas, and Patrick Cahan. 2001 · 2001
Earlier work this paper cites.
Using graded relevance assessments in IR evaluation
Jaana Kekäläinen and Kalervo Järvelin. 2002 · 2002
Earlier work this paper cites.
Retrieval evaluation with incomplete information. In
Chris Buckley and Ellen M Voorhees. 2004 · 2004
Earlier work this paper cites.
Binary and graded relevance in IR evaluations—comparison of the effects on ranking of IR systems
Jaana Kekäläinen. 2005 · 2005
Earlier work this paper cites.
Nuggeteer: Automatic nugget-based evaluation using descriptions and judgements
Gregory Marton. 2006 · 2006
Earlier work this paper cites.
In Google We Trust: Users’ Decisions on Rank, Position, and Relevance
Bing Pan, Helene Hembrooke, Thorsten Joachims, Lori Lorigo, Geri Gay, and Laura Granka. 2007 · 2007
Earlier work this paper cites.
CLEF 2008: Ad Hoc Track Overview. In
Eneko Agirre, Giorgio Maria Di Nunzio, Nicola Ferro, Thomas Mandl, and Carol Peters. 2009 · 2008
Earlier work this paper cites.
Here or there: Preference Judgments for Relevance
Ben Carterette, Paul N. Bennett, David Maxwell Chickering, and Susan T. Dumais. 2008 · 2008
Earlier work this paper cites.
Effects of inconsistent relevance judgments on information retrieval test results: A historical perspective
Tefko Saracevic. 2008 · 2008
Earlier work this paper cites.
Overview of the TREC 2009 Web Track.. In
Charles LA Clarke, Nick Craswell, and Ian Soboroff. 2009 · 2009
Earlier work this paper cites.
An introduction to information retrieval
Christopher D Manning. 2009 · 2009
Earlier work this paper cites.
A similarity measure for indefinite rankings
William Webber, Alistair Moffat, and Justin Zobel. 2010 · 2010
Earlier work this paper cites.
User Intent and Assessor Disagreement in Web Search Evaluation. In
Gabriella Kazai, Emine Yilmaz, Nick Craswell, and S.M.M. Tahaghoghi. 2013 · 2013
Earlier work this paper cites.
Ranking Retrieval Systems using Pseudo Relevance Judgments
Sri Devi Ravana, Prabha Rajagopal, and Vimala Balakrishnan. 2015 · 2015
Earlier work this paper cites.
Beyond independent relevance: methods and evaluation metrics for subtopic retrieval. In
ChengXiang Zhai, William W Cohen, and John Lafferty. 2015 · 2015
Earlier work this paper cites.
Towards Automatic Generation of Relevance Judgments for a Test Collection. In
Mireille Makary, Michael Oakes, and Fadi Yamout. 2016 · 2016
Earlier work this paper cites.
MS MARCO: A human generated machine reading comprehension dataset
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016 · 2016
Cited alongside, same era.
Using Supervised Machine Learning to Automatically Build Relevance Judgments for a Test Collection. In
Mireille Makary, Michael Oakes, Ruslan Mitkov, and Fadi Yammout. 2017 · 2017
Cited alongside, same era.
Offline evaluation without gain. In
Charles LA Clarke, Alexandra Vtyurina, and Mark D Smucker. 2020b · 2020
Cited alongside, same era.
Overview of the TREC 2020 deep learning track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, and Daniel Campos. 2021 · 2020
Cited alongside, same era.
Overview of the TREC 2019 deep learning track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Ellen M Voorhees. 2020 · 2020
Cited alongside, same era.
Query2doc: Query expansion with large language models
Liang Wang, Nan Yang, and Furu Wei. 2023a · 2023
Later among the works it cites.
Aligning large language models with human: A survey
Yufei Wang, Wanjun Zhong, Liangyou Li, Fei Mi, Xingshan Zeng, Wenyong Huang, Lifeng Shang, Xin Jiang, and Qun Liu. 2023b · 2023
Later among the works it cites.
Rank-without-GPT: Building GPT-Independent Listwise Rerankers on Open-Source Large Language Models
Xinyu Zhang, Sebastian Hofstätter, Patrick Lewis, Raphael Tang, and Jimmy Lin. 2023 · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ANTIQUE: A non-factoid question answering benchmark. In
Helia Hashemi, Mohammad Aliannejadi, Hamed Zamani, and W Bruce Croft. 2020 · 2020
Cited alongside, same era.
Good evaluation measures based on document preferences. In
Tetsuya Sakai and Zhaohao Zeng. 2020 · 2020
Cited alongside, same era.
Preference-based evaluation metrics for web image search. In
Xiaohui Xie, Jiaxin Mao, Yiqun Liu, Maarten de Rijke, Haitian Chen, Min Zhang, and Shaoping Ma. 2020 · 2020
Cited alongside, same era.
Assessing top-preferences
Charles LA Clarke, Alexandra Vtyurina, and Mark D Smucker. 2021a · 2021
Cited alongside, same era.
Assessing top-
Charles L. A. Clarke, Alexandra Vtyurina, and Mark D. Smucker. 2021b · 2021
Cited alongside, same era.
Assessing Top-
Charles L. A. Clarke, Alexandra Vtyurina, and Mark D. Smucker. 2021c · 2021
Cited alongside, same era.
Overview of the TREC 2021 deep learning track. In
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Jimmy Lin. 2022 · 2021
Cited alongside, same era.
Yutao Zhu, Huaying Yuan, Shuting Wang, Jiongnan Liu, Wenhan Liu, Chenlong Deng, Zhicheng Dou, and Ji-Rong Wen. 2023 · 2023
Later among the works it cites.
Beyond Yes and No: Improving Zero-Shot LLM Rankers via Scoring Fine-Grained Relevance Labels
Honglei Zhuang, Zhen Qin, Kai Hui, Junru Wu, Le Yan, Xuanhui Wang, and Michael Berdersky. 2023 · 2023
Later among the works it cites.
Can We Use Large Language Models to Fill Relevance Judgment Holes?
Zahra Abbasiantaeb, Chuan Meng, Leif Azzopardi, and Mohammad Aliannejadi. 2024 · 2024
Later among the works it cites.
LLMs can be Fooled into Labelling a Document as Relevant: best café near me; this paper is perfectly relevant. In
Marwah Alaofi, Paul Thomas, Falk Scholer, and Mark Sanderson. 2024b · 2024
Later among the works it cites.
Developing a framework for auditing large language models using human-in-the-loop
Maryam Amirizaniani, Jihan Yao, Adrian Lavergne, Elizabeth Snell Okada, Aman Chadha, Tanya Roosta, and Chirag Shah. 2024 · 2024
Later among the works it cites.
Fréchet Distance for Offline Evaluation of Information Retrieval Systems with Sparse Labels. In
Negar Arabzadeh and Charles Clarke. 2024a · 2024
Later among the works it cites.
A Comparison of Methods for Evaluating Generative IR
Negar Arabzadeh and Charles LA Clarke. 2024b · 2024
Later among the works it cites.
Offline Evaluation of Set-Based Text-to-Image Generation. In
Negar Arabzadeh, Fernando Diaz, and Junfeng He. 2024b · 2024
Later among the works it cites.
Assessing and Verifying Task Utility in LLM-Powered Applications. In
Negar Arabzadeh, Siqing Huo, Nikhil Mehta, Qingyun Wu, Chi Wang, Ahmed Hassan Awadallah, Charles L. A. Clarke, and Julia Kiseleva. 2024c · 2024
Later among the works it cites.
Report on The Search Futures Workshop at ECIR 2024. In
Leif Azzopardi, Charles LA Clarke, Paul Kantor, Bhaskar Mitra, Johanne R Trippas, Zhaochun Ren, Mohammad Aliannejadi, Negar Arabzadeh, Raman Chandrasekar, Maarten de Rijke, et al · 2024
Later among the works it cites.
Evaluating Relative Retrieval Effectiveness with Normalized Residual Gain. In
Amin Bigdeli, Negar Arabzadeh, Ebrahim Bagheri, and Charles LA Clarke. 2024 · 2024
Later among the works it cites.
A survey on evaluation of large language models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al · 2024
Later among the works it cites.
Steffi Chern, Ethan Chern, Graham Neubig, and Pengfei Liu. 2024 · 2024
Later among the works it cites.
LLM-based relevance assessment still can’t replace human relevance assessment
Charles L. A. Clarke and Laura Dietz. 2024 · 2024
Later among the works it cites.
Pencils down! automatic rubric-based evaluation of retrieve/generate systems. In
Naghmeh Farzi and Laura Dietz. 2024b · 2024
Later among the works it cites.
From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
Dawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi, Chengshuai Zhao, Zhen Tan, Amrita Bhattacharjee, Yuxuan Jiang, Canyu Chen, Tianhao Wu, et al · 2024
Later among the works it cites.
Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Timothy R. McIntosh, Teo Susnjak, Nalin Arachchilage, Tong Liu, Paul Watters, and Malka N. Halgamuge. 2024 · 2024
Later among the works it cites.
Query Performance Prediction using Relevance Judgments Generated by Large Language Models
Chuan Meng, Negar Arabzadeh, Arian Askari, Mohammad Aliannejadi, and Maarten de Rijke. 2024b · 2024
Later among the works it cites.
Initial nugget evaluation results for the trec 2024 rag track with the autonuggetizer framework
Ronak Pradeep, Nandan Thakur, Shivani Upadhyay, Daniel Campos, Nick Craswell, and Jimmy Lin. 2024 · 2024
Later among the works it cites.
Evaluating retrieval quality in retrieval-augmented generation. In
Alireza Salemi and Hamed Zamani. 2024 · 2024
Later among the works it cites.
A Reproducibility and Generalizability Study of Large Language Models for Query Generation. In
Moritz Staudinger, Wojciech Kusa, Florina Piroi, Aldo Lipani, and Allan Hanbury. 2024 · 2024
Later among the works it cites.
UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor
Shivani Upadhyay, Ronak Pradeep, Nandan Thakur, Nick Craswell, and Jimmy Lin. 2024c · 2024
Later among the works it cites.