Fetching the paper…
Reading the bibliography…
When asked, large language models (LLMs) like ChatGPT claim that they can assist with relevance judgments but it is not clear whether automated judgments can reliably be used in evaluations of retrieval systems.
Evaluating the Underlying Gender Bias in Contextualized Word Embeddings
Christine Basta, Marta Ruiz Costa-jussà, and Noe Casas. 2019 · 1904
Earlier work this paper cites.
Measuring Bias in Contextualized Word Representations
Keita Kurita, Nidhi Vyas, Ayush Pareek, Alan W. Black, and Yulia Tsvetkov. 2019 · 1906
Earlier work this paper cites.
The Aslib Cranfield Research Project on the Comparative Efficiency of Indexing Systems. In Aslib Proceedings , Vol. 12. MCB UP Ltd, 421–431
Cyril W Cleverdon. 1960 · 1960
Earlier work this paper cites.
Overview of the First Text REtrieval Conference (TREC-1) . NIST Special Publication, Vol. 500-207
Donna Harman. 1992 · 1992
Earlier work this paper cites.
Relevance and Information Behavior
Linda Schamber. 1994 · 1994
Earlier work this paper cites.
Preference Testing: A Comparison of Two Presentation Methods
Jennifer Windsor, Laura M Piché, and Peggy A Locke. 1994 · 1994
Earlier work this paper cites.
Evaluation of Evaluation in Information Retrieval. In SIGIR’95, Proceedings of the 18th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval. Seattle, Washington, USA, July 9-13, 1995 (Special Issue of the SIGIR Forum) . ACM Press, 138–146
Tefko Saracevic. 1995 · 1995
Earlier work this paper cites.
Relevance Reconsidered. In Proceedings of the second conference on conceptions of library and information science (CoLIS 2) . 201–218
Tefko Saracevic. 1996 · 1996
Earlier work this paper cites.
Relevance: The Whole History
Stefano Mizzaro. 1997 · 1997
Earlier work this paper cites.
Efficient Construction of Large Test Collections. In SIGIR ’98: Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, August 24-28 1998, Melbourne, Australia . ACM, 282–289
Gordon V. Cormack, Christopher R. Palmer, and Charles L. A. Clarke. 1998 · 1998
Earlier work this paper cites.
Proceedings of the First NTCIR Workshop on Research in Japanese Text Retrieval and Term Recognition . National Center for Science Information Systems (NACSIS)
Noriko Kando (Ed.). 1999 · 1999
Earlier work this paper cites.
Overview of the Eighth Text REtrieval Conference (TREC-8). In Proceedings of The Eighth Text REtrieval Conference, TREC 1999, Gaithersburg, Maryland, USA, November 17-19, 1999 (NIST Special Publication, Vol. 500-246) . National Institute of Standards and Technology (NIST)
Ellen M. Voorhees and Donna Harman. 1999 · 1999
Earlier work this paper cites.
Variations in Relevance Judgments and the Measurement of Retrieval Effectiveness
Ellen M. Voorhees. 2000 · 2000
Earlier work this paper cites.
Ranking Retrieval Systems without Relevance Judgments. In SIGIR 2001: Proceedings of the 24th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, September 9-13, 2001, New Orleans, Louisiana, USA . ACM, 66–73
Ian Soboroff, Charles K. Nicholas, and Patrick Cahan. 2001 · 2001
Earlier work this paper cites.
Using Titles and Category Names from Editor-Driven Taxonomies for Automatic Evaluation. In Proceedings of the 2003 ACM CIKM International Conference on Information and Knowledge Management, New Orleans, Louisiana, USA, November 2-8, 2003 . ACM, 17–23
Steven M. Beitzel, Eric C. Jensen, Abdur Chowdhury, and David A. Grossman. 2003 · 2003
Earlier work this paper cites.
TREC CAsT 2019: The Conversational Assistance Track Overview
Jeffrey Dalton, Chenyan Xiong, and Jamie Callan. 2020a · 2003
Earlier work this paper cites.
TREC CAsT 2019: The Conversational Assistance Track Overview
Jeffrey Dalton, Chenyan Xiong, and Jamie Callan. 2020b · 2003
Earlier work this paper cites.
Minimal Test Collections for Retrieval Evaluation. In SIGIR 2006: Proceedings of the 29th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, Seattle, Washington, USA, August 6-11, 2006 . ACM, 268–275
Ben Carterette, James Allan, and Ramesh K. Sitaraman. 2006 · 2006
Earlier work this paper cites.
Evaluating Evaluation Metrics Based on the Bootstrap. In SIGIR 2006: Proceedings of the 29th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, Seattle, Washington, USA, August 6-11, 2006 . ACM, 525–532
Tetsuya Sakai. 2006 · 2006
Earlier work this paper cites.
Charles L. A. Clarke, Alexandra Vtyurina, and Mark D. Smucker. 2020 · 2007
Earlier work this paper cites.
Relevance Assessment: Are Judges Exchangeable and Does it Matter. In Proceedings of the 31st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2008, Singapore, July 20-24, 2008 . ACM, 667–674
Peter Bailey, Nick Craswell, Ian Soboroff, Paul Thomas, Arjen P. de Vries, and Emine Yilmaz. 2008 · 2008
Earlier work this paper cites.
The YAGO-NAGA Approach to Knowledge Discovery
Gjergji Kasneci, Maya Ramanath, Fabian M. Suchanek, and Gerhard Weikum. 2008 · 2008
Earlier work this paper cites.
A Simple and Efficient Sampling Method for Estimating AP and NDCG. In Proceedings of the 31st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2008, Singapore, July 20-24, 2008 . ACM, 603–610
Emine Yilmaz, Evangelos Kanoulas, and Javed A. Aslam. 2008 · 2008
Earlier work this paper cites.
Can we get rid of TREC assessors? Using Mechanical Turk for Relevance Assessment. In Proceedings of the SIGIR 2009 Workshop on the Future of IR Evaluation , Vol. 15. 16
Omar Alonso and Stefano Mizzaro. 2009 · 2009
Earlier work this paper cites.
Explaining User Performance in Information Retrieval: Challenges to IR Evaluation. In Advances in Information Retrieval Theory, Second International Conference on the Theory of Information Retrieval, ICTIR 2009, Cambridge, UK, September 10-12, 2009, Proceedings (Lecture Notes in Computer Science, Vol. 5766) . Springer, 289–296
Kalervo Järvelin. 2009 · 2009
Earlier work this paper cites.
Estimating the Query Difficulty for Information Retrieval
David Carmel and Elad Yom-Tov. 2010 · 2010
Earlier work this paper cites.
Predicting the effectiveness of queries and retrieval systems
Claudia Hauff. 2010 · 2010
Earlier work this paper cites.
Design and Implementation of Relevance Assessments Using Crowdsourcing. In Advances in Information Retrieval - 33rd European Conference on IR Research, ECIR 2011, Dublin, Ireland, April 18-21, 2011. Proceedings (Lecture Notes in Computer Science, Vol. 6611) . Springer, 153–164
Omar Alonso and Ricardo Baeza-Yates. 2011 · 2011
Earlier work this paper cites.
Pseudo Test Collections for Learning Web Search Ranking Functions. In Proceeding of the 34th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2011, Beijing, China, July 25-29, 2011 . ACM, 1073–1082
Nima Asadi, Donald Metzler, Tamer Elsayed, and Jimmy Lin. 2011 · 2011
Earlier work this paper cites.
Repeatable and Reliable Search System Evaluation Using Crowdsourcing. In Proceeding of the 34th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2011, Beijing, China, July 25-29, 2011 . ACM, 923–932
Roi Blanco, Harry Halpin, Daniel M. Herzig, Peter Mika, Jeffrey Pound, Henry S. Thompson, and Duc Thanh Tran. 2011 · 2011
Earlier work this paper cites.
Generating Pseudo Test Collections for Learning to Rank Scientific Articles. In Information Access Evaluation. Multilinguality, Multimodality, and Visual Analytics - Third International Conference of the CLEF Initiative, CLEF 2012, Rome, Italy, September 17-20, 2012. Proceedings (Lecture Notes in Computer Science, Vol. 7488) . Springer, 42–53
Richard Berendsen, Manos Tsagkias, Maarten de Rijke, and Edgar Meij. 2012 · 2012
Earlier work this paper cites.
IR System Evaluation Using Nugget-Based Test Collections. In Proceedings of the Fifth International Conference on Web Search and Web Data Mining, WSDM 2012, Seattle, WA, USA, February 8-12, 2012 . ACM, 393–402
Virgiliu Pavlu, Shahzad Rajput, Peter B. Golbus, and Javed A. Aslam. 2012 · 2012
Cited alongside, same era.
Task partitioning effects in semi-automated human–machine system performance
PA Hancock. 2013 · 2013
Cited alongside, same era.
An Analysis of Human Factors and Label Accuracy in Crowdsourcing Relevance Judgments
Gabriella Kazai, Jaap Kamps, and Natasa Milic-Frayling. 2013 · 2013
Cited alongside, same era.
SQUARE: A Benchmark for Research on Computing Crowd Consensus. In Proceedings of the First AAAI Conference on Human Computation and Crowdsourcing, HCOMP 2013, November 7-9, 2013, Palm Springs, CA, USA . AAAI
Aashish Sheshadri and Matthew Lease. 2013 · 2013
Cited alongside, same era.
BERTScore: Evaluating Text Generation with BERT. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020 . OpenReview.net
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Later among the works it cites.
BERT-QPP: Contextualized Pre-trained Transformers for Query Performance Prediction. In CIKM ’21: The 30th ACM International Conference on Information and Knowledge Management, Virtual Event, Queensland, Australia, November 1 - 5, 2021 . ACM, 2857–2861
Negar Arabzadeh, Maryam Khodabakhsh, and Ebrahim Bagheri. 2021 · 2021
Later among the works it cites.
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?. In FAccT ’21: 2021 ACM Conference on Fairness, Accountability, and Transparency, Virtual Event / Toronto, Canada, March 3-10, 2021 . ACM, 610–623
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Later among the works it cites.
Towards Question-Answering as an Automatic Metric for Evaluating the Content Quality of a Summary
Daniel Deutsch, Tania Bedrax-Weiss, and Dan Roth. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gaya K. Jayasinghe, William Webber, Mark Sanderson, and J. Shane Culpepper. 2014 · 2014
Cited alongside, same era.
Evaluating Answer Passages Using Summarization Measures. In The 37th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’14, Gold Coast , QLD, Australia - July 06 - 11, 2014 . ACM, 963–966
Mostafa Keikha, Jae Hyun Park, and W. Bruce Croft. 2014 · 2014
Cited alongside, same era.
Automatic Gloss Finding for a Knowledge Base using Ontological Constraints. In Proceedings of the Eighth ACM International Conference on Web Search and Data Mining, WSDM 2015, Shanghai, China, February 2-6, 2015 . ACM, 369–378
Bhavana Bharat Dalvi, Einat Minkov, Partha Pratim Talukdar, and William W. Cohen. 2015 · 2015
Cited alongside, same era.
SIGIR 2014: Workshop on Gathering Efficient Assessments of Relevance (GEAR)
Martin Halvey, Robert Villa, and Paul D. Clough. 2015 · 2015
Cited alongside, same era.
Shared Control is the Sharp End of Cooperation: Towards a Common Framework of Joint Action, Shared Control and Human Machine Cooperation
Frank Flemisch, David Abbink, Makoto Itoh, Marie-Pierre Pacaux-Lemoine, and Gina Weßel. 2016 · 2016
Cited alongside, same era.
WikiReading: A Novel Large-scale Language Understanding Task over Wikipedia
Daniel Hewlett, Alexandre Lacoste, Llion Jones, Illia Polosukhin, Andrew Fandrianto, Jay Han, Matthew Kelcey, and David Berthelot. 2016 · 2016
Cited alongside, same era.
Crowdsourcing Relevance Assessments: The Unexpected Benefits of Limiting the Time to Judge. In Proceedings of the Fourth AAAI Conference on Human Computation and Crowdsourcing, HCOMP 2016, 30 October - 3 November, 2016, Austin, Texas, USA . AAAI Press, 129–138
Eddy Maddalena, Marco Basaldella, Dario De Nart, Dante Degl’Innocenti, Stefano Mizzaro, and Gianluca Demartini. 2016 · 2016
Cited alongside, same era.
On the Impact of Domain Expertise on Query Formulation, Relevance Assessment and Retrieval Performance in Clinical Settings
Lynda Tamine and Cecile Chouquet. 2017 · 2016
Cited alongside, same era.
Later among the works it cites.
System Effect Estimation by Sharding: A Comparison Between ANOVA Approaches to Detect Significant Differences. In Advances in Information Retrieval - 43rd European Conference on IR Research, ECIR 2021, Virtual Event, March 28 - April 1, 2021, Proceedings, Part II (Lecture Notes in Computer Science, Vol. 12657) . Springer, 33–46
Guglielmo Faggioli and Nicola Ferro. 2021 · 2021
Later among the works it cites.
WikiAsp: A Dataset for Multi-domain Aspect-based Summarization
Hiroaki Hayashi, Prashant Budania, Peng Wang, Chris Ackerson, Raj Neervannan, and Graham Neubig. 2021 · 2021
Later among the works it cites.
EXAM: How to Evaluate Retrieve-and-Generate Systems for Users Who Do Not (Yet) Know What They Want. In Proceedings of the Second International Conference on Design of Experimental Search & Information REtrieval Systems, Padova, Italy, September 15-18, 2021 (CEUR Workshop Proceedings, Vol. 2950) . CEUR-WS.org, 136–146
David P. Sander and Laura Dietz. 2021 · 2021
Later among the works it cites.
Overview of TREC 2021. In 30th Text REtrieval Conference . Gaithersburg, Maryland
Ian Soboroff. 2021 · 2021
Later among the works it cites.
Unsupervised Question Clarity Prediction through Retrieved Item Coherency. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management, Atlanta, GA, USA, October 17-21, 2022 . ACM, 3811–3816
Negar Arabzadeh, Mahsa Seifikar, and Charles L. A. Clarke. 2022 · 2022
Later among the works it cites.
Groupwise Query Performance Prediction with BERT. In Advances in Information Retrieval - 44th European Conference on IR Research, ECIR 2022, Stavanger, Norway, April 10-14, 2022, Proceedings, Part II (Lecture Notes in Computer Science, Vol. 13186) . Springer, 64–74
Xiaoyang Chen, Ben He, and Le Sun. 2022 · 2022
Later among the works it cites.
Faithful Reasoning Using Large Language Models
Antonia Creswell and Murray Shanahan. 2022 · 2022
Later among the works it cites.
A ’Pointwise-Query, Listwise-Document’ Based Query Performance Prediction Approach. In Proceedings of 45th international ACM SIGIR conference research development in information retrieval . 2148––2153
Suchana Datta, Sean MacAvaney, Debasis Ganguly, and Derek Greene. 2022 · 2022
Later among the works it cites.
Wikimarks: Harvesting Relevance Benchmarks from Wikipedia. In SIGIR ’22: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, Madrid, Spain, July 11 - 15, 2022 . ACM, 3003–3012
Laura Dietz, Shubham Chatterjee, Connor Lennox, Sumanta Kashyapi, Pooja Oza, and Ben Gamari. 2022 · 2022
Later among the works it cites.
Noise-Reduction for Automatically Transferred Relevance Judgments. In Experimental IR Meets Multilinguality, Multimodality, and Interaction - 13th International Conference of the CLEF Association, CLEF 2022, Bologna, Italy, September 5-8, 2022, Proceedings (Lecture Notes in Computer Science, Vol. 13390) . Springer, 48–61
Maik Fröbe, Christopher Akiki, Martin Potthast, and Matthias Hagen. 2022 · 2022
Later among the works it cites.
FIRE ’22: Proceedings of the 14th Annual Meeting of the Forum for Information Retrieval Evaluation (Kolkata, India). Association for Computing Machinery, New York, NY, USA
Debasis Ganguly, Surupendu Gangopadhyay, Mandar Mitra, and Prasenjit Majumder (Eds.). 2022 · 2022
Later among the works it cites.
Active Learning on Pre-trained Language Model with Task-Independent Triplet Loss. In Thirty-Fourth Conference on Innovative Applications of Artificial Intelligence, IAAI 2022, The Twelveth Symposium on Educational Advances in Artificial Intelligence, EAAI 2022 Virtual Event, February 22 - March 1, 2022 . AAAI Press, 11276–11284
Seungmin Seo, Donghyun Kim, Youbin Ahn, and Kyong-Ho Lee. 2022 · 2022
Later among the works it cites.
An Information-theoretic Approach to Prompt Engineering Without Ground Truth Labels. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022 . Association for Computational Linguistics, 819–862
Taylor Sorensen, Joshua Robinson, Christopher Michael Rytting, Alexander Glenn Shaw, Kyle Jeffrey Rogers, Alexia Pauline Delorey, Mahmoud Khalil, Nancy Fulda, and David Wingate. 2022 · 2022
Later among the works it cites.
Leveraging Similar Users for Personalized Language Modeling with Limited Data. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022 . Association for Computational Linguistics, 1742–1752
Charles Welch, Chenxi Gu, Jonathan K. Kummerfeld, Verónica Pérez-Rosas, and Rada Mihalcea. 2022 · 2022
Later among the works it cites.
On The Role of Human and Machine Metadata in Relevance Judgment Tasks
Jiechen Xu, Lei Han, Shazia Sadiq, and Gianluca Demartini. 2023 · 2022
Later among the works it cites.
Multimodal Knowledge Alignment with Reinforcement Learning
Youngjae Yu, Jiwan Chung, Heeseung Yun, Jack Hessel, Jae Sung Park, Ximing Lu, Prithviraj Ammanabrolu, Rowan Zellers, Ronan Le Bras, Gunhee Kim, and Yejin Choi. 2022 · 2022
Later among the works it cites.
Large Language Models Are Human-Level Prompt Engineers
Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. 2022 · 2022
Later among the works it cites.
Christine Bauer, Ben Carterette, Nicola Ferro, and Norbert Fuhr. 2023 · 2023
Closest in time.
A Geometric Framework for Query Performance Prediction in Conversational Search. In Proceedings of 46th international ACM SIGIR Conference on Research & Development in Information Retrieval, SIGIR 2023 July 23–27, 2023, Taipei, Taiwan . ACM
Guglielmo Faggioli, Nicola Ferro, Cristina Muntean, Raffaele Perego, and Nicola Tonellotto. 2023 · 2023
Closest in time.
ExaRanker: Explanation-Augmented Neural Ranker
Fernando Ferraretto, Thiago Laitz, Roberto de Alencar Lotufo, and Rodrigo Nogueira. 2023 · 2023
Closest in time.
Bootstrapped nDCG Estimation in the Presence of Unjudged Documents. In Advances in Information Retrieval - 45th European Conference on Information Retrieval, ECIR 2023, Dublin, Ireland, April 2-6, 2023, Proceedings, Part I (Lecture Notes in Computer Science, Vol. 13980) . Springer, 313–329
Maik Fröbe, Lukas Gienapp, Martin Potthast, and Matthias Hagen. 2023 · 2023
Closest in time.
ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023 · 2023
Closest in time.
Faithful Chain-of-Thought Reasoning
Qing Lyu, Shreya Havaldar, Adam Stein, Li Zhang, Delip Rao, Eric Wong, Marianna Apidianaki, and Chris Callison-Burch. 2023 · 2023
Closest in time.
One-Shot Labeling for Automatic Relevance Estimation
Sean MacAvaney and Luca Soldaini. 2023 · 2023
Closest in time.
Can ChatGPT Reproduce Human-Generated Labels? A Study of Social Computing Tasks
Yiming Zhu, Peixian Zhang, Ehsan ul Haq, Pan Hui, and Gareth Tyson. 2023 · 2023
Closest in time.
CLEF 2000 - Overview of Results. In Cross-Language Information Retrieval and Evaluation, Workshop of Cross-Language Evaluation Forum, CLEF 2000, Lisbon, Portugal, September 21-22, 2000, Revised Papers (Lecture Notes in Computer Science, Vol. 2069) . Springer, 89–101
Martin Braschler. 2000 · 2069
Closest in time.