Fetching the paper…
Reading the bibliography…
The application of large language models to provide relevance assessments presents exciting opportunities to advance information retrieval, natural language processing, and beyond, but to date many unknowns remain.
Large Language Models Can Accurately Predict Searcher Preferences. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2024) . Washington, D.C., 1930–1940
Paul Thomas, Seth Spielman, Nick Craswell, and Bhaskar Mitra. 2024 · 1940
Earlier work this paper cites.
Relevance Assessments and Retrieval System Evaluation
Michael E. Lesk and Gerard Salton. 1968 · 1968
Earlier work this paper cites.
Variations in Relevance Judgments and the Measurement of Retrieval Effectiveness. In Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 1998) . Melbourne, Australia, 315–323
Ellen M. Voorhees. 1998 · 1998
Earlier work this paper cites.
How Reliable Are the Results of Large-Scale Information Retrieval Experiments?. In Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 1998) . Melbourne, Australia, 307–314
Justin Zobel. 1998 · 1998
Earlier work this paper cites.
Evaluating Evaluation Measure Stability. In Proceedings of the 23rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2000) . Athens, Greece, 33–40
Chris Buckley and Ellen M. Voorhees. 2000 · 2000
Earlier work this paper cites.
Passage-Based Query Refinement: (MultiText Experiments for TREC-6)
Gordon V. Cormack, Charles L.A. Clarke, Christopher R. Palmer, and Samuel S.L. To. 2000 · 2000
Earlier work this paper cites.
Retrieval Evaluation with Incomplete Information. In Proceedings of the 27th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2004) . Sheffield, United Kingdom, 25–32
Chris Buckley and Ellen M. Voorhees. 2004 · 2004
Earlier work this paper cites.
Accurately Interpreting Clickthrough Data as Implicit Feedback. In Proceedings of the 28th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2005) . Salvador, Brazil, 154–161
Thorsten Joachims, Laura Granka, Bing Pang, Helene Hembrooke, and Geri Gay. 2005 · 2005
Earlier work this paper cites.
Information Retrieval System Evaluation: Effort, Sensitivity, and Reliability. In Proceedings of the 28th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2005) . Salvador, Brazil, 162–169
Mark Sanderson and Justin Zobel. 2005 · 2005
Earlier work this paper cites.
Information Retrieval Evaluation
Donna Harman. 2011 · 2011
Earlier work this paper cites.
Community-Based Bayesian Aggregation Models for Crowdsourcing. In Proceedings of the 23rd International World Wide Web Conference (WWW 2014) . Seoul, South Korea, 155–164
Matteo Venanzi, John Guiver, Gabriella Kazai, Pushmeet Kohli, and Milad Shokouhi. 2014 · 2014
Earlier work this paper cites.
MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosenberg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang. 2018 · 2018
Cited alongside, same era.
Constitutional AI: Harmlessness from AI Feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Kamile Lukosuite, Liane Lovitt, Michael Sellitto, Nelson Elhage, Nicholas Schiefer, Noemi Mercado, Nova DasSarma, Robert Lasenby, Robin Larson, Sam Ringer, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Tamera Lanham, Timothy Telleen-Lawton, Tom Conerly, Tom Henighan, Tristan Hume, Samuel R. Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan. 2022 · 2022
Cited alongside, same era.
InPars: Data Augmentation for Information Retrieval using Large Language Models. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2022) . Madrid, Spain, 2387–2392
Luiz Bonifacio, Hugo Abonizio, Marzieh Fadaee, and Rodrigo Nogueira. 2022 · 2022
ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023 · 2023
Later among the works it cites.
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Later among the works it cites.
A Comparison of Methods for Evaluating Generative IR
Negar Arabzadeh and Charles L. A. Clarke. 2024 · 2024
Closest in time.
Investigating Data Contamination in Modern Benchmarks for Large Language Models
Chunyuan Deng, Yilun Zhao, Xiangru Tang, Mark Gerstein, and Arman Cohan. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Overview of the TREC 2022 Deep Learning Track. In Proceedings of the Thirty-First Text REtrieval Conference (TREC 2022) . Gaithersburg, Maryland
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, Jimmy Lin, Ellen M. Voorhees, and Ian Soboroff. 2022 · 2022
Cited alongside, same era.
Language Models Trained on Media Diets Can Predict Public Opinion
Eric Chu, Jacob Andreas, Stephen Ansolabehere, and Deb Roy. 2023 · 2023
Cited alongside, same era.
LM vs LM: Detecting Factual Errors via Cross Examination
Roi Cohen, May Hamri, Mor Geva, and Amir Globerson. 2023 · 2023
Cited alongside, same era.
Overview of the TREC 2023 Deep Learning Track. In Proceedings of the Thirty-Second Text REtrieval Conference (TREC 2023) . Gaithersburg, Maryland
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Hossein A. Rahmani, Daniel Campos, Jimmy Lin, Ellen M. Voorhees, and Ian Soboroff. 2023 · 2023
Cited alongside, same era.
Promptagator: Few-shot Dense Retrieval From 8 Examples. In Proceedings of the 11th International Conference on Learning Representations (ICLR 2023)
Zhuyun Dai, Vincent Y. Zhao, Ji Ma, Yi Luan, Jianmo Ni, Jing Lu, Anton Bakalov, Kelvin Guu, Keith B. Hall, and Ming-Wei Chang. 2023 · 2023
Cited alongside, same era.
Perspectives on Large Language Models for Relevance Judgment. In Proceedings of the 2023 ACM SIGIR International Conference on Theory of Information Retrieval . Taipei, Taiwan
Guglielmo Faggioli, Laura Dietz, Charles L. A. Clarke, Gianluca Demartini, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Benno Stein, and Henning Wachsmuth. 2023 · 2023
Cited alongside, same era.
GPTScore: Evaluate as You Desire
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023 · 2023
Cited alongside, same era.
Human-like Summarization Evaluation with ChatGPT
Mingqi Gao, Jie Ruan, Renliang Sun, Xunjian Yin, Shiping Yang, and Xiaojun Wan. 2023 · 2023
Cited alongside, same era.
Ricardo Dominguez-Olmedo, Moritz Hardt, and Celestine Mendler-Dünner. 2024 · 2024
Closest in time.
AI-Augmented Surveys: Leveraging Large Language Models and Surveys for Opinion Prediction
Junsol Kim and Byungkyu Lee. 2024 · 2024
Closest in time.
Ragnarök: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented Generation Track
Ronak Pradeep, Nandan Thakur, Sahel Sharifymoghaddam, Eric Zhang, Ryan Nguyen, Daniel Campos, Nick Craswell, and Jimmy Lin. 2024 · 2024
Closest in time.
LLMJudge: LLMs for Relevance Judgments
Hossein A. Rahmani, Emine Yilmaz, Nick Craswell, Bhaskar Mitra, Paul Thomas, Charles L. A. Clarke, Mohammad Aliannejadi, Clemencia Siro, and Guglielmo Faggioli. 2024 · 2024
Closest in time.
Researchy Questions: A Dataset of Multi-Perspective, Decompositional Questions for LLM Web Agents
Corby Rosset, Ho-Lam Chung, Guanghui Qin, Ethan C. Chau, Zhuo Feng, Ahmed Awadallah, Jennifer Neville, and Nikhil Rao. 2024 · 2024
Closest in time.
Don’t Use LLMs to Make Relevance Judgments
Ian Soboroff. 2024 · 2024
Closest in time.
UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor
Shivani Upadhyay, Ronak Pradeep, Nandan Thakur, Nick Craswell, and Jimmy Lin. 2024 · 2024
Closest in time.