Fetching the paper…
Reading the bibliography…
Relevance labels, which indicate whether a search result is valuable to a searcher, are key to evaluating and optimising search systems.
Problems of monetary management: The UK experience
Charles A E Goodhart. 1975 · 1975
Earlier work this paper cites.
OHSUMED: An interactive retrieval evaluation and new large test collection for research. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval . 192–201
William Hersh, Chris Buckley, TJ Leone, and David Hickam. 1994 · 1994
Earlier work this paper cites.
The ‘awful’ idea of accountability: Inscribing people into the measurement of objects
Keith Hoskin. 1996 · 1996
Earlier work this paper cites.
Efficient construction of large test collections. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval . 282–289
Gordon V Cormack, Christopher R Palmer, and Charles L A Clarke. 1998 · 1998
Earlier work this paper cites.
Variations in relevance judgments and the measurement of retrieval effectiveness. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval . 315–323
Ellen M Voorhees. 1998 · 1998
Earlier work this paper cites.
Goodhart’s law: Its origins, meaning and implications for monetary policy
K. Alec Chrystal and Paul D. Mizen. 2001 · 2001
Earlier work this paper cites.
A taxonomy of web search. In ACM Sigir forum , Vol. 36. ACM New York, NY, USA, 3–10
Andrei Broder. 2002 · 2002
Earlier work this paper cites.
Overview of the TREC 2004 Robust Retrieval Track. In Proceedings of the Text REtrieval Conference
Ellen M Voorhees. 2004 · 2004
Earlier work this paper cites.
A reference collection for web spam
Carlos Castillo, Debora Donato, Luca Becchetti, Paolo Boldi, Stefano Leonardi, Massimo Santini, and Sebastiano Vigna. 2006 · 2006
Earlier work this paper cites.
Relevance assessment: Are judges exchangeable and does it matter?. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval . 667–674
Peter Bailey, Nick Craswell, Ian Soboroff, Paul Thomas, Arjen P. de Vries, and Emine Yilmaz. 2008 · 2008
Earlier work this paper cites.
Here or there: Preference judgments for relevance. In Proceedings of the European Conference on Information Retrieval . 16–27
Ben Carterette, Paul N Bennett, David Maxwell Chickering, and Susan T Dumais. 2008 · 2008
Earlier work this paper cites.
Effects of inconsistent relevance judgments on information retrieval test results: A historical perspective
Tefko Saracevic. 2008 · 2008
Earlier work this paper cites.
Learning to rank for information retrieval
Tie-Yan Liu. 2009 · 2009
Earlier work this paper cites.
Test collection based evaluation of information retrieval systems
Mark Sanderson. 2010 · 2010
Earlier work this paper cites.
A similarity measure for indefinite rankings
William Webber, Alistair Moffat, and Justin Zobel. 2010 · 2010
Earlier work this paper cites.
Examining the limits of crowdsourcing for relevance assessment
Paul Clough, Mark Sanderson, Jiayu Tang, Tim Gollins, and Amy Warner. 2013 · 2013
Earlier work this paper cites.
The effect of threshold priming and need for cognition on relevance calibration and assessment. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval . 623–632
Falk Scholer, Diane Kelly, Wan-Ching Wu, Hanseul S. Lee, and William Webber. 2013 · 2013
Earlier work this paper cites.
Discrimination in online ad delivery
Latanya Sweeney. 2013 · 2013
Earlier work this paper cites.
Understanding user behavior through log data and analysis
Susan Dumais, Robin Jeffries, , Daniel M. Russell, Diane Tang, and Jaime Teevan. 2014 · 2014
Earlier work this paper cites.
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Cited alongside, same era.
Man is to computer programmer as woman is to homemaker? Debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016 · 2016
Cited alongside, same era.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017 · 2017
Cited alongside, same era.
Gauging the quality of relevance assessments using inter-rater agreement. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval
Tadele T. Damessie, Taho P. Nghiem, Falk Scholer, and J. Shane Culpepper. 2017 · 2017
Cited alongside, same era.
Algorithms of oppression
Safiya Umoja Noble. 2018 · 2018
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2022 · 2022
Later among the works it cites.
Sustainable AI: Environmental implications, challenges and opportunities
Carole-Jean Wu, Ramya Raghavendra, Udit Gupta, Bilge Acun, Newsha Ardalani, Kiwan Maeng, Gloria Chang, Fiona Aga, Jinshi Huang, Charles Bai, et al · 2022
Later among the works it cites.
TEMPERA: Test-time prompt editing via reinforcement learning
Tianjun Zhang, Xuezhi Wang, Denny Zhou, Dale Schuurmans, and Joseph E. Gonzalez. 2022 · 2022
Later among the works it cites.
Large language models are human-level prompt engineers
Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Good, neutral or bad news classification. In Proceedings of the Third International Workshop on Recent Trends in News Information Retrieval . 9–14
Aashish Agarwal, Ankita Mandal, Matthias Schaffeld, Fangzheng Ji, Jhiao Zhan, Yiqi Sun, and Ahmet Aker. 2019 · 2019
Cited alongside, same era.
Lipstick on a pig: Debiasing methods cover up systematic gender biases in word embeddings but do not remove them. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . 609–614
Hila Gonen and Yoav Goldberg. 2019 · 2019
Cited alongside, same era.
Language (technology) is power: A critical survey of “bias”’ in NLP. In Proceedings of the Annual Meeting of the Association for Computational Linguistics . 5454–5476
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020 · 2020
Cited alongside, same era.
Do people and neural nets pay attention to the same words: studying eye-tracking data for non-factoid QA evaluation. In Proceedings of the ACM International Conference on Information and Knowledge Management . 85–94
Valeria Bolotova, Vladislav Blinov, Yukun Zheng, W Bruce Croft, Falk Scholer, and Mark Sanderson. 2020 · 2020
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Cited alongside, same era.
Carbon emissions and large neural network training
David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean. 2021 · 2021
Cited alongside, same era.
Meysam Alizadeh, Maël Kubli, Zeynab Samei, Shirin Dehghani, Juan Diego Bermeo, Maria Korobeynikovo, and Fabrizio Gilardi. 2023 · 2023
Closest in time.
Can large language models be an alternative to human evaluations?. In Proceedings of the Annual Meeting of the Association for Computational Linguistics . 15607–15631
Cheng-Han Chiang and Hung-yi Lee. 2023 · 2023
Closest in time.
HMC: A spectrum of human–machine-collaborative relevance judgment frameworks
Charles L A Clarke, Gianluca Demartini, Laura Dietz, Guglielmo Faggioli, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Ian Soboroff, Benno Stein, and Henning Wachsmuth. 2023 · 2023
Closest in time.
Perspectives on large language models for relevance judgment
Guglielmo Faggioli, Laura Dietz, Charles Clarke, Gianluca Demartini, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Benno Stein, and Henning Wachsmuth. 2023 · 2023
Closest in time.
ChatGPT outperforms crowd-workers for text-annotation tasks
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023 · 2023
Closest in time.
General Guidelines
Google LLC. 2022 · 2023
Closest in time.
Collect, measure, repeat: Reliability factors for responsible AI data collection
Oana Inel, Tim Draws, and Lora Aroyo. 2023 · 2023
Closest in time.
State of GPT
Andrej Karpathy. 2023 · 2023
Closest in time.
G-EVAL: NLG evaluation using GPT-4 with better human alignment
Yang Liu, Dan Iter, Yichong xu amd Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Automatic prompt optimization with “gradient descent” and beam search
Reid Pryzant, Dan Iter, Jerry Li, Yin Tat Lee, Chenguang Zhu, and Michael Zeng. 2023 · 2023
Closest in time.
Petter Törnberg. 2023 · 2023
Closest in time.
How far can camels go? Exploring the state of instruction tuning on open resources
Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, David Wadden, Kelsey MacMillan, Noah A. Smith, Iz Beltagy, and Hannaneh Hajishirzi. 2023 · 2023
Closest in time.
Large language models as optimisers
Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V. Le, Denny Zhou, and Xinyun Chen. 2023 · 2023
Closest in time.
Local self-attention over long text for efficient document retrieval. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval . 2021–2024
Sebastian Hofstätter, Hamed Zamani, Bhaskar Mitra, Nick Craswell, and Allan Hanbury. 2020 · 2024
Closest in time.