Fetching the paper…
Reading the bibliography…
Test collections are information-retrieval tools that allow researchers to quickly and easily evaluate ranking algorithms.
A Sequential Algorithm for Training Text Classifiers
David D. Lewis and William A. Gale. 1994 · 1994
Earlier work this paper cites.
Efficient construction of large test collections. In Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (Melbourne, Australia) (SIGIR ’98) . Association for Computing Machinery, New York, NY, USA, 282–289
Gordon V. Cormack, Christopher R. Palmer, and Charles L. A. Clarke. 1998 · 1998
Earlier work this paper cites.
Variations in relevance judgments and the measurement of retrieval effectiveness. In Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (Melbourne, Australia) (SIGIR ’98) . Association for Computing Machinery, New York, NY, USA, 315–323
Ellen M. Voorhees. 1998 · 1998
Earlier work this paper cites.
Overview of the Eighth Text REtrieval Conference (TREC-8). In Proceedings of the Eighth Text REtrieval Conference (TREC-8) (NIST Special Publication, 500-246) , E.M. Voorhees and D.K. Harman (Eds.). 1–24
Ellen M. Voorhees and Donna Harman. 2000 · 2000
Earlier work this paper cites.
Evaluating Content Selection in Summarization: The Pyramid Method. In Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics: HLT-NAACL 2004 . Association for Computational Linguistics, Boston, Massachusetts, USA, 145–152
Ani Nenkova and Rebecca Passonneau. 2004 · 2004
Earlier work this paper cites.
The TREC robust retrieval track
Ellen M. Voorhees. 2005 · 2005
Earlier work this paper cites.
A statistical method for system evaluation using incomplete judgments. In Proceedings of the 29th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (Seattle, Washington, USA) (SIGIR ’06) . Association for Computing Machinery, New York, NY, USA, 541–548
Javed A. Aslam, Virgil Pavlu, and Emine Yilmaz. 2006 · 2006
Earlier work this paper cites.
Inferring document relevance from incomplete information. In Proceedings of the Sixteenth ACM Conference on Conference on Information and Knowledge Management (Lisbon, Portugal) (CIKM ’07) . Association for Computing Machinery, New York, NY, USA, 633–642
Javed A. Aslam and Emine Yilmaz. 2007 · 2007
Earlier work this paper cites.
Relevance assessment: are judges exchangeable and does it matter. In Proceedings of the 31st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (Singapore, Singapore) (SIGIR ’08) . Association for Computing Machinery, New York, NY, USA, 667–674
Peter Bailey, Nick Craswell, Ian Soboroff, Paul Thomas, Arjen P. de Vries, and Emine Yilmaz. 2008 · 2008
Earlier work this paper cites.
Machine learning for information retrieval: TREC 2009 Web, Relevance Feedback and Legal Tracks. In The Eighteenth Text REtrieval Conference (TREC 2009)
G. V. Cormack and M. Mojdeh. 2009 · 2009
Earlier work this paper cites.
Design and implementation of relevance assessments using crowdsourcing. In Proceedings of the 33rd European Conference on Advances in Information Retrieval (Dublin, Ireland) (ECIR’11) . Springer-Verlag, Berlin, Heidelberg, 153–164
Omar Alonso and Ricardo Baeza-Yates. 2011 · 2011
Earlier work this paper cites.
Examining the Limits of Crowdsourcing for Relevance Assessment
Paul Clough, Mark Sanderson, Jiayu Tang, Tim Gollins, and Amy Warner. 2013 · 2012
Earlier work this paper cites.
Quality through flow and immersion: gamifying crowdsourced relevance assessments. In Proceedings of the 35th International ACM SIGIR Conference on Research and Development in Information Retrieval (Portland, Oregon, USA) (SIGIR ’12) . Association for Computing Machinery, New York, NY, USA, 871–880
Carsten Eickhoff, Christopher G. Harris, Arjen P. de Vries, and Padmini Srinivasan. 2012 · 2012
Earlier work this paper cites.
Evaluation of machine-learning protocols for technology-assisted review in electronic discovery. In Proceedings of the 37th International ACM SIGIR Conference on Research & Development in Information Retrieval (Gold Coast, Queensland, Australia) (SIGIR ’14) . Association for Computing Machinery, New York, NY, USA, 153–162
Gordon V. Cormack and Maura R. Grossman. 2014 · 2014
Earlier work this paper cites.
Scalability of Continuous Active Learning for Reliable High-Recall Text Classification. In Proceedings of the 25th ACM International on Conference on Information and Knowledge Management (Indianapolis, Indiana, USA) (CIKM ’16) . Association for Computing Machinery, New York, NY, USA, 1039–1048
Gordon V. Cormack and Maura R. Grossman. 2016 · 2016
Cited alongside, same era.
Feeling lucky? multi-armed bandits for ordering judgements in pooling-based evaluation. In Proceedings of the 31st Annual ACM Symposium on Applied Computing (Pisa, Italy) (SAC ’16) . Association for Computing Machinery, New York, NY, USA, 1027–1034
David E. Losada, Javier Parapar, and Álvaro Barreiro. 2016 · 2016
Cited alongside, same era.
Active Sampling for Large-scale Information Retrieval Evaluation. In Proceedings of the 2017 ACM on Conference on Information and Knowledge Management (CIKM ’17) . ACM
Dan Li and Evangelos Kanoulas. 2017 · 2017
Cited alongside, same era.
Jointly Minimizing the Expected Costs of Review for Responsiveness and Privilege in E-Discovery
Douglas W. Oard, Fabrizio Sebastiani, and Jyothi K. Vinjumur. 2018 · 2018
Debasis Ganguly and Emine Yilmaz. 2023 · 2023
Later among the works it cites.
One-Shot Labeling for Automatic Relevance Estimation. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’23) . ACM
Sean MacAvaney and Luca Soldaini. 2023 · 2023
Later among the works it cites.
Stopping Methods for Technology-assisted Reviews Based on Point Processes
Mark Stevenson and Reem Bin-Hezam. 2023 · 2023
Later among the works it cites.
Open-Domain Dialogue Quality Evaluation: Deriving Nugget-level Scores from Turn-level Scores. In Proceedings of the Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region (SIGIR-AP ’23) . ACM, 40–45
Rikiya Takehi, Akihisa Watanabe, and Tetsuya Sakai. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On Building Fair and Reusable Test Collections using Bandit Techniques. In Proceedings of the 27th ACM International Conference on Information and Knowledge Management (Torino, Italy) (CIKM ’18) . Association for Computing Machinery, New York, NY, USA, 407–416
Ellen M. Voorhees. 2018 · 2018
Cited alongside, same era.
Humans Optional? Automatic Large-Scale Test Collections for Entity, Passage, and Entity-Passage Retrieval
L. Dietz and Jeff Dalton. 2020 · 2020
Cited alongside, same era.
Efficient Test Collection Construction via Active Learning. In Proceedings of the 2020 ACM SIGIR on International Conference on Theory of Information Retrieval (Virtual Event, Norway) (ICTIR ’20) . Association for Computing Machinery, New York, NY, USA, 177–184
Md Mustafizur Rahman, Mucahid Kutlu, Tamer Elsayed, and Matthew Lease. 2020 · 2020
Cited alongside, same era.
Searching for Scientific Evidence in a Pandemic: An Overview of TREC-COVID
Kirk Roberts, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, Kyle Lo, Ian Soboroff, Ellen Voorhees, Lucy Lu Wang, and William R Hersh. 2021 · 2021
Cited alongside, same era.
BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021 · 2021
Cited alongside, same era.
Wikimarks: Harvesting Relevance Benchmarks from Wikipedia. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (Madrid, Spain) (SIGIR ’22) . Association for Computing Machinery, New York, NY, USA, 3003–3012
Laura Dietz, Shubham Chatterjee, Connor Lennox, Sumanta Kashyapi, Pooja Oza, and Ben Gamari. 2022 · 2022
Cited alongside, same era.
The Crowd is Made of People: Observations from Large-Scale Crowd Labelling. In Proceedings of the 2022 Conference on Human Information Interaction and Retrieval (Regensburg, Germany) (CHIIR ’22) . Association for Computing Machinery, New York, NY, USA, 25–35
Paul Thomas, Gabriella Kazai, Ryen White, and Nick Craswell. 2022 · 2022
Cited alongside, same era.
Too Many Relevants: Whither Cranfield Test Collections?. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (Madrid, Spain) (SIGIR ’22) . Association for Computing Machinery, New York, NY, USA, 2970–2980
Ellen M. Voorhees, Nick Craswell, and Jimmy Lin. 2022 · 2022
Cited alongside, same era.
Zahra Abbasiantaeb, Chuan Meng, Leif Azzopardi, and Mohammad Aliannejadi. 2024 · 2024
Closest in time.
LLMs can be Fooled into Labelling a Document as Relevant: best cafe near me; this paper is perfectly relevant. In Proceedings of the 2024 Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region (Tokyo, Japan) (SIGIR-AP 2024) . Association for Computing Machinery, New York, NY, USA, 32–41
Marwah Alaofi, Paul Thomas, Falk Scholer, and Mark Sanderson. 2024 · 2024
Closest in time.
LLM-based relevance assessment still can’t replace human relevance assessment
Charles L. A. Clarke and Laura Dietz. 2024 · 2024
Closest in time.
Large Language Models for Relevance Judgment in Product Search
Navid Mehrdad, Hrushikesh Mohapatra, Mossaab Bagdouri, Prijith Chandran, Alessandro Magnani, Xunfan Cai, Ajit Puthenputhussery, Sachin Yadav, Tony Lee, ChengXiang Zhai, and Ciya Liao. 2024 · 2024
Closest in time.
Query Performance Prediction using Relevance Judgments Generated by Large Language Models
Chuan Meng, Negar Arabzadeh, Arian Askari, Mohammad Aliannejadi, and Maarten de Rijke. 2024 · 2024
Closest in time.
Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting
Zhen Qin, Rolf Jagerman, Kai Hui, Honglei Zhuang, Junru Wu, Le Yan, Jiaming Shen, Tianqi Liu, Jialu Liu, Donald Metzler, Xuanhui Wang, and Michael Bendersky. 2024 · 2024
Closest in time.
Synthetic Test Collections for Retrieval Evaluation
Hossein A. Rahmani, Nick Craswell, Emine Yilmaz, Bhaskar Mitra, and Daniel Campos. 2024 · 2024
Closest in time.
Don’t Use LLMs to Make Relevance Judgments
Ian Soboroff. 2024 · 2024
Closest in time.
Large language models can accurately predict searcher preferences
Paul Thomas, Seth Spielman, Nick Craswell, and Bhaskar Mitra. 2024 · 2024
Closest in time.