Fetching the paper…
Reading the bibliography…
The traditional evaluation of information retrieval (IR) systems is generally very costly as it requires manual relevance annotation from human experts.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Relevance assessments and retrieval system evaluation
Michael E Lesk and Gerard Salton. 1968 · 1968
Earlier work this paper cites.
Document clustering: An evaluation of some experiments with the Cranfield 1400 collection
Cornelis Joost Van Rijsbergen and W Bruce Croft. 1975 · 1975
Earlier work this paper cites.
Better bootstrap confidence intervals
Bradley Efron. 1987 · 1987
Earlier work this paper cites.
A review of bootstrap confidence intervals
Thomas J Diciccio and Joseph P Romano. 1988 · 1988
Earlier work this paper cites.
The relationship between recall and precision
Michael Buckland and Fredric Gey. 1994 · 1994
Earlier work this paper cites.
Bootstrap confidence intervals
Thomas J DiCiccio and Bradley Efron. 1996 · 1996
Earlier work this paper cites.
Overview of IR tasks at the first NTCIR workshop. In Proceedings of the first NTCIR workshop on research in Japanese text retrieval and term recognition . 11–44
Noriko Kando, Kazuko Kuriyama, Toshihiko Nozue, Koji Eguchi, Hiroyuki Kato, and Souichiro Hidaka. 1999 · 1999
Earlier work this paper cites.
Cumulated gain-based evaluation of IR techniques
Kalervo Järvelin and Jaana Kekäläinen. 2002 · 2002
Earlier work this paper cites.
Using graded relevance assessments in IR evaluation
Jaana Kekäläinen and Kalervo Järvelin. 2002 · 2002
Earlier work this paper cites.
Overview of the TREC 2003 robust retrieval track.. In Trec . 69–77
Ellen M Voorhees et al · 2003
Earlier work this paper cites.
The TREC test collections
Donna K Harman. 2005 · 2005
Earlier work this paper cites.
Agreement, the f-measure, and reliability in information retrieval
George Hripcsak and Adam S Rothschild. 2005 · 2005
Earlier work this paper cites.
Information retrieval system evaluation: effort, sensitivity, and reliability. In Proceedings of the 28th annual international ACM SIGIR conference on Research and development in information retrieval . 162–169
Mark Sanderson and Justin Zobel. 2005 · 2005
Earlier work this paper cites.
TREC: Experiment and evaluation in information retrieval . Vol. 63
Ellen M Voorhees, Donna K Harman, et al · 2005
Earlier work this paper cites.
A statistical method for system evaluation using incomplete judgments. In Proceedings of the 29th annual international ACM SIGIR conference on Research and development in information retrieval . 541–548
Javed A Aslam, Virgil Pavlu, and Emine Yilmaz. 2006 · 2006
Earlier work this paper cites.
Minimal test collections for retrieval evaluation. In Proceedings of the 29th annual international ACM SIGIR conference on Research and development in information retrieval . 268–275
Ben Carterette, James Allan, and Ramesh Sitaraman. 2006 · 2006
Earlier work this paper cites.
Statistical precision of information retrieval evaluation. In Proceedings of the 29th annual international ACM SIGIR conference on Research and development in information retrieval . 533–540
Gordon V Cormack and Thomas R Lynam. 2006 · 2006
Earlier work this paper cites.
A comparison of statistical significance tests for information retrieval evaluation. In Proceedings of the sixteenth ACM conference on Conference on information and knowledge management . 623–632
Mark D Smucker, James Allan, and Ben Carterette. 2007 · 2007
Earlier work this paper cites.
Relevance assessment: are judges exchangeable and does it matter. In Proceedings of the 31st annual international ACM SIGIR conference on Research and development in information retrieval . 667–674
Peter Bailey, Nick Craswell, Ian Soboroff, Paul Thomas, Arjen P de Vries, and Emine Yilmaz. 2008 · 2008
Earlier work this paper cites.
Inductive conformal prediction: Theory and application to neural networks
Harris Papadopoulos. 2008 · 2008
Earlier work this paper cites.
A simple and efficient sampling method for estimating AP and NDCG. In Proceedings of the 31st annual international ACM SIGIR conference on Research and development in information retrieval . 603–610
Emine Yilmaz, Evangelos Kanoulas, and Javed A Aslam. 2008 · 2008
Earlier work this paper cites.
LETOR: A benchmark collection for research on learning to rank for information retrieval
Tao Qin, Tie-Yan Liu, Jun Xu, and Hang Li. 2010 · 2010
Earlier work this paper cites.
Test collection based evaluation of information retrieval systems
Mark Sanderson et al · 2010
Earlier work this paper cites.
Measurement in information retrieval evaluation
William Edward Webber. 2010 · 2010
Earlier work this paper cites.
Yahoo! learning to rank challenge overview. In Proceedings of the learning to rank challenge . PMLR, 1–24
Olivier Chapelle and Yi Chang. 2011 · 2011
Earlier work this paper cites.
Information retrieval evaluation
Donna Harman. 2011 · 2011
Cited alongside, same era.
Bootstrap
Tim Hesterberg. 2011 · 2011
Cited alongside, same era.
Introducing LETOR 4.0 datasets
Tao Qin and Tie-Yan Liu. 2013 · 2013
Cited alongside, same era.
Approximate recall confidence intervals
William Webber. 2013 · 2013
Cited alongside, same era.
Conformal prediction for reliable machine learning: theory, adaptations and applications
Vineeth Balasubramanian, Shen-Shyang Ho, and Vladimir Vovk. 2014 · 2014
Cited alongside, same era.
Statistical reform in information retrieval?. In ACM SIGIR Forum , Vol. 48. ACM New York, NY, USA, 3–12
Tetsuya Sakai. 2014 · 2014
Cited alongside, same era.
TREC deep learning track: Reusable test collections in the large data regime. In Proceedings of the 44th international ACM SIGIR conference on research and development in information retrieval . 2369–2375
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, Ellen M Voorhees, and Ian Soboroff. 2021 · 2021
Later among the works it cites.
Beir: A heterogenous benchmark for zero-shot evaluation of information retrieval models
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021 · 2021
Later among the works it cites.
TREC-COVID: constructing a pandemic information retrieval test collection. In ACM SIGIR Forum , Vol. 54. ACM New York, NY, USA, 1–12
Ellen Voorhees, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, William R Hersh, Kyle Lo, Kirk Roberts, Ian Soboroff, and Lucy Lu Wang. 2021 · 2021
Later among the works it cites.
Anastasios N Angelopoulos, Stephen Bates, Adam Fisch, Lihua Lei, and Tal Schuster. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multilingual representations for low resource speech recognition and keyword search. In 2015 IEEE workshop on automatic speech recognition and understanding (ASRU) . IEEE, 259–266
Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran, Abhinav Sethy, Kartik Audhkhasi, Xiaodong Cui, Ellen Kislal, Lidia Mangu, Markus Nussbaum-Thom, Michael Picheny, et al · 2015
Cited alongside, same era.
An overview of the BIOASQ large-scale biomedical semantic indexing and question answering competition
George Tsatsaronis, Georgios Balikas, Prodromos Malakasiotis, Ioannis Partalas, Matthias Zschunke, Michael R Alvers, Dirk Weissenborn, Anastasia Krithara, Sergios Petridis, Dimitris Polychronopoulos, et al · 2015
Cited alongside, same era.
Predicting relevance based on assessor disagreement: analysis and practical applications for search evaluation
Thomas Demeester, Robin Aly, Djoerd Hiemstra, Dong Nguyen, and Chris Develder. 2016 · 2016
Cited alongside, same era.
MS MARCO: A human generated machine reading comprehension dataset
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016 · 2016
Cited alongside, same era.
Model-assisted survey estimation with modern prediction techniques
F Jay Breidt and Jean D Opsomer. 2017 · 2017
Cited alongside, same era.
IR evaluation methods for retrieving highly relevant documents. In ACM SIGIR Forum , Vol. 51. ACM New York, NY, USA, 243–250
Kalervo Järvelin and Jaana Kekäläinen. 2017 · 2017
Cited alongside, same era.
Inpars: Unsupervised dataset generation for information retrieval. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval . 2387–2392
Luiz Bonifacio, Hugo Abonizio, Marzieh Fadaee, and Rodrigo Nogueira. 2022 · 2022
Later among the works it cites.
The Istella22 Dataset: Bridging Traditional and Neural Learning to Rank Evaluation. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval . 3099–3107
Domenico Dato, Sean MacAvaney, Franco Maria Nardini, Raffaele Perego, and Nicola Tonellotto. 2022 · 2022
Later among the works it cites.
Generative artificial intelligence: Trends and prospects
Mladan Jovanovic and Mark Campbell. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Unifying language learning paradigms
Yi Tay, Mostafa Dehghani, Vinh Q Tran, Xavier Garcia, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Neil Houlsby, and Donald Metzler. 2022 · 2022
Later among the works it cites.
Meysam Alizadeh, Maël Kubli, Zeynab Samei, Shirin Dehghani, Juan Diego Bermeo, Maria Korobeynikova, and Fabrizio Gilardi. 2023 · 2023
Later among the works it cites.
Anastasios N Angelopoulos, Stephen Bates, Clara Fannjiang, Michael I Jordan, and Tijana Zrnic. 2023 · 2023
Later among the works it cites.
4.2 HMC: A Spectrum of Human–Machine-Collaborative Relevance Judgment Frameworks
Charles LA Clarke, Gianluca Demartini, Laura Dietz, Guglielmo Faggioli, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Ian Soboroff, et al · 2023
Later among the works it cites.
Perspectives on large language models for relevance judgment. In Proceedings of the 2023 ACM SIGIR International Conference on Theory of Information Retrieval . 39–50
Guglielmo Faggioli, Laura Dietz, Charles LA Clarke, Gianluca Demartini, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Benno Stein, et al · 2023
Later among the works it cites.
Chatgpt outperforms crowd-workers for text-annotation tasks
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023 · 2023
Later among the works it cites.
Improving Cross-lingual Information Retrieval on Low-Resource Languages via Optimal Transport Distillation. In Proceedings of the Sixteenth ACM International Conference on Web Search and Data Mining . 1048–1056
Zhiqi Huang, Puxuan Yu, and James Allan. 2023 · 2023
Later among the works it cites.
Generative artificial intelligence and its applications in materials science: Current situation and future perspectives
Yue Liu, Zhengwei Yang, Zhenyao Yu, Zitu Liu, Dahui Liu, Hailong Lin, Mingqing Li, Shuchang Ma, Maxim Avdeev, and Siqi Shi. 2023 · 2023
Later among the works it cites.
The flan collection: Designing data and methods for effective instruction tuning
Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V Le, Barret Zoph, Jason Wei, et al · 2023
Later among the works it cites.
One-Shot Labeling for Automatic Relevance Estimation
Sean MacAvaney and Luca Soldaini. 2023 · 2023
Later among the works it cites.
Collaborating with ChatGPT: Considering the implications of generative artificial intelligence for journalism and media education
John V Pavlik. 2023 · 2023
Later among the works it cites.
Large language models are effective text rankers with pairwise ranking prompting
Zhen Qin, Rolf Jagerman, Kai Hui, Honglei Zhuang, Junru Wu, Jiaming Shen, Tianqi Liu, Jialu Liu, Donald Metzler, Xuanhui Wang, et al · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Later among the works it cites.
Large language models can accurately predict searcher preferences
Paul Thomas, Seth Spielman, Nick Craswell, and Bhaskar Mitra. 2023 · 2023
Later among the works it cites.
Petter Törnberg. 2023 · 2023
Later among the works it cites.
Beyond yes and no: Improving zero-shot llm rankers via scoring fine-grained relevance labels
Honglei Zhuang, Zhen Qin, Kai Hui, Junru Wu, Le Yan, Xuanhui Wang, and Michael Berdersky. 2023 · 2023
Later among the works it cites.
Cross-Language Information Retrieval and Evaluation: Workshop of Cross-Language Evaluation Forum, CLEF 2000, Lisbon, Portugal, September 21-22, 2000, Revised Papers . Vol. 2069
Carol Peters. 2001 · 2069
Closest in time.