Fetching the paper…
Reading the bibliography…
Copious amounts of relevance judgments are necessary for the effective training and accurate evaluation of retrieval systems.
Overview of the TREC 2019 Deep Learning Track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Ellen M. Voorhees. 2020 · 2003
Earlier work this paper cites.
Overview of the TREC 2020 Deep Learning Track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, and Daniel Campos. 2021 · 2020
Earlier work this paper cites.
Humans Optional? Automatic Large-Scale Test Collections for Entity, Passage, and Entity-Passage Retrieval
Laura Dietz and Jeff Dalton. 2020 · 2020
Earlier work this paper cites.
Overview of the TREC 2021 Deep Learning Track. In Text REtrieval Conference (TREC) . NIST, TREC
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Jimmy Lin. 2022 · 2021
Earlier work this paper cites.
Overview of the TREC 2022 Deep Learning Track. In Text REtrieval Conference (TREC) . NIST, TREC
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, Jimmy Lin, Ellen M. Voorhees, and Ian Soboroff. 2023 · 2022
Earlier work this paper cites.
Wikimarks: Harvesting Relevance Benchmarks from Wikipedia. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (Madrid, Spain). Association for Computing Machinery, 3003–3012
Laura Dietz, Shubham Chatterjee, Connor Lennox, Sumanta Kashyapi, Pooja Oza, and Ben Gamari. 2022 · 2022
Earlier work this paper cites.
Meysam Alizadeh, Maël Kubli, Zeynab Samei, Shirin Dehghani, Juan Diego Bermeo, Maria Korobeynikovo, and Fabrizio Gilardi. 2023 · 2023
Earlier work this paper cites.
Can Large Language Models be an Alternative to Human Evaluations? 15607–15631
Cheng-Han Chiang and Hung-yi Lee. 2023 · 2023
Earlier work this paper cites.
Overview of the TREC 2023 Deep Learning Track. In Text REtrieval Conference (TREC) . NIST, TREC
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Hossein A. Rahmani, Daniel Campos, Jimmy Lin, Ellen M. Voorhees, and Ian Soboroff. 2024 · 2023
Earlier work this paper cites.
Perspectives on Large Language Models for Relevance Judgment
Guglielmo Faggioli, Laura Dietz, Charles Clarke, Gianluca Demartini, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Benno Stein, and Henning Wachsmuth. 2023 · 2023
Cited alongside, same era.
ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023 · 2023
Cited alongside, same era.
Large Language Models are Zero-Shot Reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2023 · 2023
Cited alongside, same era.
G-EVAL: NLG Evaluation Using GPT-4 with Better Human Alignment
Yang Liu, Dan Iter, Yichong xu amd Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Cited alongside, same era.
Open-Source LLMs for Text Annotation: A Practical Guide for Model Setting and Fine-Tuning
Meysam Alizadeh, Maël Kubli, Zeynab Samei, Shirin Dehghani, Mohammadmasiha Zahedivafa, Juan Diego Bermeo, Maria Korobeynikova, and Fabrizio Gilardi. 2024 · 2024
Closest in time.
AnnoLLM: Making Large Language Models to Be Better Crowdsourced Annotators
Xingwei He, Zhenghao Lin, Yeyun Gong, A-Long Jin, Hang Zhang, Chen Lin, Jian Jiao, Siu Ming Yiu, Nan Duan, and Weizhu Chen. 2024 · 2024
Closest in time.
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2024
Closest in time.
Large Language Models for Data Annotation: A Survey
Zhen Tan, Alimohammad Beigi, Song Wang, Ruocheng Guo, Amrita Bhattacharjee, Bohan Jiang, Mansooreh Karami, Jundong Li, Lu Cheng, and Huan Liu. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
One-Shot Labeling for Automatic Relevance Estimation. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’23)
Sean MacAvaney and Luca Soldaini. 2023 · 2023
Cited alongside, same era.
Can GPT-4 Support Analysis of Textual Data in Tasks Requiring Highly Specialized Domain Expertise?
Jaromir Savelka, Kevin D Ashley, Morgan A Gray, Hannes Westermann, and Huihui Xu. 2023 · 2023
Cited alongside, same era.
Petter Törnberg. 2023 · 2023
Cited alongside, same era.
Can ChatGPT Reproduce Human-Generated Labels? a Study of Social Computing Tasks
Yiming Zhu, Peixian Zhang, Ehsan-Ul Haq, Pan Hui, and Gareth Tyson. 2023 · 2023
Cited alongside, same era.
Gemini: A Family of Highly Capable Multimodal Models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, Anja Hauth, Katie Millican, David Silver, et al · 2024
Closest in time.
Large Language Models Can Accurately Predict Searcher Preferences. In 2024 International ACM SIGIR Conference on Research and Development in Information Retrieval . ACM
Paul Thomas, Seth Spielman, Nick Craswell, and Bhaskar Mitra. 2024 · 2024
Closest in time.
Best Practices for Text Annotation with Large Language Models
Petter Törnberg. 2024 · 2024
Closest in time.
LLMs Can Patch Up Missing Relevance Judgments in Evaluation
Shivani Upadhyay, Ehsan Kamalloo, and Jimmy Lin. 2024 · 2024
Closest in time.