Fetching the paper…
Reading the bibliography…
The first edition of the workshop on Large Language Model for Evaluation in Information Retrieval (LLM4Eval 2024) took place in July 2024, co-located with the ACM SIGIR Conference 2024 in the USA (SIGIR 2024).
Ms marco: A human generated machine reading comprehension dataset
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng · 2016
Earlier work this paper cites.
Evaluating cross-modal generative models using retrieval task
Shivangi Bithel and Srikanta Bedathur · 2023
Earlier work this paper cites.
Overview of the trec 2023 deep learning track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Hossein A. Rahmani, Daniel Campos, Jimmy Lin, Ellen M. Voorhees, and Ian Soboroff · 2023
Earlier work this paper cites.
One-shot labeling for automatic relevance estimation
Sean MacAvaney and Luca Soldaini · 2023
Earlier work this paper cites.
Large language models can accurately predict searcher preferences
Paul Thomas, Seth Spielman, Nick Craswell, and Bhaskar Mitra · 2023
Earlier work this paper cites.
Can we use large language models to fill relevance judgment holes?
Zahra Abbasiantaeb, Chuan Meng, Leif Azzopardi, and Mohammad Aliannejadi · 2024
Earlier work this paper cites.
Bhashithe Abeysinghe and Ruhan Circi · 2024
Earlier work this paper cites.
Evaluating the retrieval component in llm-based question answering systems
Ashkan Alinejad, Krtin Kumar, and Ali Vahdat · 2024
Earlier work this paper cites.
A comparison of methods for evaluating generative ir
Negar Arabzadeh and Charles LA Clarke · 2024
Cited alongside, same era.
Exploring large language models for relevance judgments in tetun
Gabriel de Jesus and Sérgio Nunes · 2024
Cited alongside, same era.
Exam++: Llm-based answerability metrics for ir evaluation
Naghmeh Farzi and Laura Dietz · 2024
Cited alongside, same era.
A novel evaluation framework for image2text generation, 2024
Jia-Hong Huang, Hongyi Zhu, Yixian Shen, Stevan Rudinac, Alessio M. Pacces, and Evangelos Kanoulas · 2024
Cited alongside, same era.
Using llms to investigate correlations of conversational follow-up queries with user satisfaction
Hyunwoo Kim, Yoonseo Choi, Taehyun Yang, Honggu Lee, Chaneon Park, Yongju Lee, Jin Young Kim, and Juho Kim · 2024
Query performance prediction using relevance judgments generated by large language models
Chuan Meng, Negar Arabzadeh, Arian Askari, Mohammad Aliannejadi, and Maarten de Rijke · 2024
Closest in time.
Reliable confidence intervals for information retrieval evaluation using generative ai
Harrie Oosterhuis, Rolf Jagerman, Zhen Qin, Xuanhui Wang, and Michael Bendersky · 2024
Closest in time.
Evaluating rag-fusion with ragelo: an automated elo-based framework
Zackary Rackauckas, Arthur Câmara, and Jakub Zavrel · 2024
Closest in time.
Clemencia Siro, Mohammad Aliannejadi, and Maarten de Rijke · 2024
Closest in time.
Followir: Evaluating and teaching information retrieval models to follow instructions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bhawesh Kumar, Jonathan Amar, Eric Yang, Nan Li, and Yugang Jia · 2024
Cited alongside, same era.
On the evaluation of machine-generated reports
James Mayfield, Eugene Yang, Dawn Lawrie, Sean MacAvaney, Paul McNamee, Douglas W Oard, Luca Soldaini, Ian Soboroff, Orion Weller, Efsun Kayi, et al · 2024
Cited alongside, same era.
Large language models for relevance judgment in product search
Navid Mehrdad, Hrushikesh Mohapatra, Mossaab Bagdouri, Prijith Chandran, Alessandro Magnani, Xunfan Cai, Ajit Puthenputhussery, Sachin Yadav, Tony Lee, ChengXiang Zhai, et al · 2024
Cited alongside, same era.
Synthetic test collections for retrieval evaluation
Hossein A Rahmani, Nick Craswell, Emine Yilmaz, Bhaskar Mitra, and Daniel Campos
Cited in the paper.
Llm4eval: Large language model for evaluation in ir
Hossein A Rahmani, Clemencia Siro, Mohammad Aliannejadi, Nick Craswell, Charles LA Clarke, Guglielmo Faggioli, Bhaskar Mitra, Paul Thomas, and Emine Yilmaz
Cited in the paper.
Orion Weller, Benjamin Chang, Sean MacAvaney, Kyle Lo, Arman Cohan, Benjamin Van Durme, Dawn Lawrie, and Luca Soldaini · 2024
Closest in time.
Jheng-Hong Yang and Jimmy Lin · 2024
Closest in time.
Weijia Zhang, Mohammad Aliannejadi, Yifei Yuan, Jiahuan Pei, Jia-Hong Huang, and Evangelos Kanoulas · 2024
Closest in time.