Fetching the paper…
Reading the bibliography…
Making the relevance judgments for a TREC-style test collection can be complex and expensive.
Report on the need for and provision of an ”ideal” information retrieval test collection
K. Sparck Jones and C. van Rijsbergen · 1975
Earlier work this paper cites.
Introduction to Modern Information Retrieval
Gerard Salton and Michael J. McGill · 1983
Earlier work this paper cites.
Indexing by latent semantic analysis
Scott Deerwester, Susan T. Dumais, George W. Furnas, Thomas K. Landauer, and Richard A. Harshman · 1990
Earlier work this paper cites.
Automatic query expansion using SMART: TREC 3
Chris Buckley, Gerard Salton, James Allan, and Amit Singhal · 1994
Earlier work this paper cites.
Overview of the fourth Text REtrieval Conference (TREC-4)
Donna Harman · 1995
Earlier work this paper cites.
The Cranfield tests on index language devices
C. W. Cleverdon · 1997
Earlier work this paper cites.
Variations in relevance judgments and the measurement of retrieval effectiveness
Ellen M. Voorhees · 1998
Earlier work this paper cites.
How reliable are the results of large-scale information retrieval experiments?
Justin Zobel · 1998
Earlier work this paper cites.
Ranking retrieval systems without relevance judgments
Ian Soboroff, Charles Nicholas, and Patrick Cahan · 2001
Earlier work this paper cites.
The philosophy of information retrieval evaluation
Ellen M. Voorhees · 2001
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
On the effectiveness of evaluating retrieval systems in the absence of relevance judgments
Javed A. Aslam and Robert Savell · 2003
Earlier work this paper cites.
Overview of the TREC 2004 web track
Nick Craswell and David Hawking · 2004
Earlier work this paper cites.
UMass at TREC 2004: Novelty and HARD
Nasreen Abdul Jaleel, James Allan, W. Bruce Croft, Fernando Diaz, Leah S. Larkey, Xiaoyan Li, Mark D. Smucker, and Courtney Wade · 2004
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Novelty detection: The TREC experience
Ian Soboroff and Donna Harman · 2005
Cited alongside, same era.
TREC: Experiment and Evaluation in Information Retrieval
Ellen M. Voorhees and Donna K. Harman, editors · 2005
Cited alongside, same era.
Bias and the limits of pooling for large collections
Chris Buckley, Darrin Dimmick, Ian Soboroff, and Ellen Voorhees · 2007
Cited alongside, same era.
Reliable information retrieval evaluation with incomplete and biased judgements
Stefan Büttcher, Charles L. A. Clarke, Peter C. K. Yeung, and Ian Soboroff · 2007
Cited alongside, same era.
A new rank correlation coefficient for information retrieval
Emine Yilmaz, Javed A. Aslam, and Stephen Robertson · 2008
Cited alongside, same era.
Overview of the TREC 2011 entity track
Krisztian Balog, Pavel Serdyukov, and Arjen P. de Vries · 2011
Cited alongside, same era.
Humans rely more on algorithms than social influence as a task becomes more difficult
Eric Bogert, Aaron Schechter, and Richard T. Watson · 2021
Later among the works it cites.
TREC CAsT 2022: Going beyond user ask and system retrieve with initiative and response generation
Paul Owoicho, Jeff Dalton, Mohammad Aliannejadi, Leif Azzopardi, Johanne R. Trippas, and Svitlana Vakulenko · 2022
Later among the works it cites.
Can old TREC collections reliably evaluate modern neural retrieval models?
Ellen M. Voorhees, Ian Soboroff, and Jimmy Lin · 2022
Later among the works it cites.
TREC ikat 2023: The interactive knowledge assistance track overview
Mohammad Aliannejadi, Zahra Abbasiantaeb, Shubham Chatterjee, Jeffery Dalton, and Leif Azzopardi · 2023
Later among the works it cites.
Report on the Dagstuhl seminar on frontiers of information access experimentation for research and education
Christine Bauer, Ben Carterette, Nicola Ferro, Norbert Fuhr, Joeran Beel, Timo Breuer, Charles L. A. Clarke, Anita Crescenzi, Gianluca Demartini, Giorgio Maria Di Nunzio, Laura Dietz, Guglielmo Faggioli, Bruce Ferwerda, Maik Fröbe, Matthias Hagen, Allan Hanbury, Claudia Hauff, Dietmar Jannach, Noriko Kando, Evangelos Kanoulas, Bart P. Knijnenburg, Udo Kruschwitz, Meijie Li, Maria Maistro, Lien Michiels, Andrea Papenmeier, Martin Potthast, Paolo Rosso, Alan Said, Philipp Schaer, Christin Seifert, Damiano Spina, Benno Stein, Nava Tintarev, Julián Urbano, Henning Wachsmuth, Martijn C. Willemsen, and Justin Zobel · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Google effects on memory: Cognitive consequences of having information at our fingertips
Betsy Sparrow, Jenny Liu, and Daniel M. Wegner · 2011
Cited alongside, same era.
Constructing test collections by inferring document relevance via extracted relevant information
Shahzad Rajput, Matthew Ekstrand-Abueg, Virgil Pavlu, and Javed A. Aslam · 2012
Cited alongside, same era.
Overview of the TREC 2014 session track
Ben Carterette, Evangelos Kanoulas, Mark M. Hall, and Paul D. Clough · 2014
Cited alongside, same era.
Evaluating stream filtering for entity profile updates in TREC 2012, 2013, and 2014
John R. Frank, Max Kleiman-Weiner, Daniel A. Roberts, Ellen M. Voorhees, and Ian Soboroff · 2014
Cited alongside, same era.
MS MARCO: A human generated machine reading comprehension dataset
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng · 2016
Cited alongside, same era.
Algorithm appreciation: People prefer algorithmic to human judgment
Jennifer M. Logg, Julia A. Minson, and Don A. Moore · 2018
Cited alongside, same era.
Later among the works it cites.
Exploring the use of large language models for reference-free text quality evaluation: An empirical study
Yi Chen, Rui Wang, Haiyun Jiang, Shuming Shi, and Ruifeng Xu · 2023
Later among the works it cites.
LLM-eval: Unified multi-dimensional automatic evaluation for open-domain conversations with large language models
Yen-Ting Lin and Yun-Nung Chen · 2023
Later among the works it cites.
One-shot labeling for automatic relevance estimation
Sean MacAvaney and Luca Soldaini · 2023
Later among the works it cites.
What’s the meaning of superhuman performance in today’s NLU?
Simone Tedeschi, Johan Bos, Thierry Declerck, Jan Hajič, Daniel Hershcovich, Eduard Hovy, Alexander Koller, Simon Krek, Steven Schockaert, Rico Sennrich, Ekaterina Shutova, and Roberto Navigli · 2023
Later among the works it cites.
LLMs can be fooled into labelling a document as relevant: best café near me; this paper is perfectly relevant
Marwah Alaofi, Paul Thomas, Falk Scholer, and Mark Sanderson · 2024
Closest in time.
Who determines what is relevant? Humans or AI? Why not both?
Guglielmo Faggioli, Laura Dietz, Charles L. A. Clarke, Gianluca Demartini, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Benno Stein, and Henning Wachsmuth · 2024
Closest in time.
GPTScore: Evaluate as you desire
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu · 2024
Closest in time.
The curious decline of linguistic diversity: Training language models on synthetic text, 2024
Yanzhu Guo, Guokan Shang, Michalis Vazirgiannis, and Chloé Clavel · 2024
Closest in time.
Large language models can accurately predict searcher preferences
Paul Thomas, Seth Spielman, Nick Craswell, and Bhaskar Mitra · 2024
Closest in time.