Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are increasingly integral to information retrieval (IR), powering ranking, evaluation, and AI-assisted content creation.
Rodrigo Nogueira and Kyunghyun Cho. 2019 · 1901
Earlier work this paper cites.
Structured Pruning of a BERT-based Question Answering Model
J. S. McCarley, Rishav Chakravarti, and Avirup Sil. 2019 · 1910
Earlier work this paper cites.
Large Language Models can Accurately Predict Searcher Preferences. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’24) . 1930–1940
Paul Thomas, Seth Spielman, Nick Craswell, and Bhaskar Mitra. 2024 · 1940
Earlier work this paper cites.
Relevance reconsidered. In Proceedings of the second conference on conceptions of library and information science (CoLIS 2) . 201–218
Tefko Saracevic. 1996 · 1996
Earlier work this paper cites.
Overview of the TREC 2019 Deep Learning track. In Proceedings of the Twenty-Eighth Text REtrieval Conference (TREC ’19)
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Ellen M. Voorhees. 2019 · 2019
Earlier work this paper cites.
Text Summarization with Pretrained Encoders. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) . 3730–3740
Yang Liu and Mirella Lapata. 2019 · 2019
Earlier work this paper cites.
MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP ’19) . 563–578
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019 · 2019
Earlier work this paper cites.
Overview of the TREC 2020 Deep Learning Track. In Proceedings of the Twenty-Ninth Text REtrieval Conference (TREC ’20)
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, and Daniel Campos. 2020 · 2020
Earlier work this paper cites.
Document ranking with a pretrained sequence-to-sequence model. In Findings of the Association for Computational Linguistics: EMNLP 2020 . 708–718
Rodrigo Nogueira, Zhiying Jiang, and Jimmy Lin. 2020 · 2020
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Earlier work this paper cites.
BERTScore: Evaluating Text Generation with BERT. In 8th International Conference on Learning Representations (ICLR ’20)
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Earlier work this paper cites.
Incorporating BERT into Neural Machine Translation. In International Conference on Learning Representations (ICLR ’20)
Jinhua Zhu, Yingce Xia, Lijun Wu, Di He, Tao Qin, Wengang Zhou, Houqiang Li, and Tieyan Liu. 2020 · 2020
Earlier work this paper cites.
Towards understanding and mitigating social biases in language models. In International Conference on Machine Learning (ICML ’21) . 6565–6576
Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2021 · 2021
Earlier work this paper cites.
Pyserini: A Python Toolkit for Reproducible Information Retrieval Research with Sparse and Dense Representations. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’21) . 2356–2362
Jimmy Lin, Xueguang Ma, Sheng-Chieh Lin, Jheng-Hong Yang, Ronak Pradeep, and Rodrigo Nogueira. 2021 · 2021
Earlier work this paper cites.
The Expando-Mono-Duo Design Pattern for Text Ranking with Pretrained Sequence-to-Sequence Models
Ronak Pradeep, Rodrigo Nogueira, and Jimmy Lin. 2021 · 2021
Earlier work this paper cites.
InPars: Unsupervised Dataset Generation for Information Retrieval. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’22) . 2387–2392
Luiz Bonifacio, Hugo Abonizio, Marzieh Fadaee, and Rodrigo Nogueira. 2022 · 2022
Earlier work this paper cites.
Improving Passage Retrieval with Zero-Shot Question Generation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP ’22) . 3781–3797
Devendra Sachan, Mike Lewis, Mandar Joshi, Armen Aghajanyan, Wen-tau Yih, Joelle Pineau, and Luke Zettlemoyer. 2022 · 2022
Earlier work this paper cites.
Transformer memory as a differentiable search index. In Proceedings of the 36th International Conference on Neural Information Processing Systems (NeurIPS ’22) . 21831–21843
Yi Tay, Vinh Tran, Mostafa Dehghani, Jianmo Ni, Dara Bahri, Harsh Mehta, Zhen Qin, Kai Hui, Zhe Zhao, Jai Gupta, et al · 2022
Earlier work this paper cites.
Promptagator: Few-shot Dense Retrieval From 8 Examples. In The Eleventh International Conference on Learning Representations (ICLR ’23)
Zhuyun Dai, Vincent Y Zhao, Ji Ma, Yi Luan, Jianmo Ni, Jing Lu, Anton Bakalov, Kelvin Guu, Keith Hall, and Ming-Wei Chang. 2023 · 2023
Earlier work this paper cites.
PaRaDe: Passage Ranking using Demonstrations with LLMs. In Findings of the Association for Computational Linguistics: EMNLP 2023 (EMNLP ’23) . 14242–14252
Andrew Drozdov, Honglei Zhuang, Zhuyun Dai, Zhen Qin, Razieh Rahimi, Xuanhui Wang, Dana Alon, Mohit Iyyer, Andrew McCallum, Donald Metzler, and Kai Hui. 2023 · 2023
Earlier work this paper cites.
Perspectives on Large Language Models for Relevance Judgment. In Proceedings of the 2023 ACM SIGIR International Conference on Theory of Information Retrieval (ICTIR ’23) . 39–50
Guglielmo Faggioli, Laura Dietz, Charles L. A. Clarke, Gianluca Demartini, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Benno Stein, and Henning Wachsmuth. 2023 · 2023
Cited alongside, same era.
Gemini: A family of highly capable multimodal models
Gemini Team Google. 2023 · 2023
Cited alongside, same era.
ChatGPT outperforms crowd workers for text-annotation tasks
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023 · 2023
Cited alongside, same era.
Holistic Evaluation of Language Models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Alexander Cosgrove, Christopher D Manning, Christopher Re, Diana Acosta-Navas, Drew Arad Hudson, Eric Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue WANG, Keshav Santhanam, Laurel Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan Kim, Neel Guha, Niladri S. Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Andrew Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda. 2023 · 2023
Lost in the middle: How language models use long contexts
Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024a · 2024
Later among the works it cites.
LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores. In Findings of the Association for Computational Linguistics: ACL 2024 . 12688–12701
Yiqi Liu, Nafise Moosavi, and Chenghua Lin. 2024b · 2024
Later among the works it cites.
Large Language Models for Relevance Judgment in Product Search
Navid Mehrdad, Hrushikesh Mohapatra, Mossaab Bagdouri, Prijith Chandran, Alessandro Magnani, Xunfan Cai, Ajit Puthenputhussery, Sachin Yadav, Tony Lee, ChengXiang Zhai, and Ciya Liao. 2024 · 2024
Later among the works it cites.
Reliable confidence intervals for information retrieval evaluation using generative ai. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD ’24) . 2307–2317
Harrie Oosterhuis, Rolf Jagerman, Zhen Qin, Xuanhui Wang, and Michael Bendersky. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP ’23) . 2511–2522
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Cited alongside, same era.
Zero-Shot Listwise Document Reranking with a Large Language Model
Xueguang Ma, Xinyu Zhang, Ronak Pradeep, and Jimmy Lin. 2023 · 2023
Cited alongside, same era.
One-Shot Labeling for Automatic Relevance Estimation. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’23) . 2230–2235
Sean MacAvaney and Luca Soldaini. 2023 · 2023
Cited alongside, same era.
Biases in large language models: origins, inventory, and discussion
Roberto Navigli, Simone Conia, and Björn Ross. 2023 · 2023
Cited alongside, same era.
Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP ’23) . 14918–14937
Weiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang, Pengjie Ren, Zhumin Chen, Dawei Yin, and Zhaochun Ren. 2023 · 2023
Cited alongside, same era.
How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NeurIPS ’23) . 74764–74786
Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Chandu, David Wadden, Kelsey MacMillan, Noah A Smith, Iz Beltagy, and Hannaneh Hajishirzi. 2023 · 2023
Cited alongside, same era.
Judging LLM-as-a-judge with MT-bench and Chatbot Arena. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NeurIPS ’23)
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Cited alongside, same era.
RankT5: Fine-Tuning T5 for Text Ranking with Ranking Losses. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’23) . 2308–2313
Honglei Zhuang, Zhen Qin, Rolf Jagerman, Kai Hui, Ji Ma, Jing Lu, Jianmo Ni, Xuanhui Wang, and Michael Bendersky. 2023 · 2023
Cited alongside, same era.
LLM Evaluators Recognize and Favor Their Own Generations. In The Thirty-eighth Annual Conference on Neural Information Processing Systems (NeurIPS ’24)
Arjun Panickssery, Samuel R. Bowman, and Shi Feng. 2024 · 2024
Later among the works it cites.
Analyzing Adversarial Attacks on Sequence-to-Sequence Relevance Models. In Proceedings of the 46th European Conference on Information Retrieval (ECIR ’24) . 286–302
Andrew Parry, Maik Fröbe, Sean MacAvaney, Martin Potthast, and Matthias Hagen. 2024 · 2024
Later among the works it cites.
Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting. In Findings of the Association for Computational Linguistics: NAACL 2024 . 1504–1518
Zhen Qin, Rolf Jagerman, Kai Hui, Honglei Zhuang, Junru Wu, Le Yan, Jiaming Shen, Tianqi Liu, Jialu Liu, Donald Metzler, Xuanhui Wang, and Michael Bendersky. 2024 · 2024
Later among the works it cites.
Hossein A. Rahmani, Clemencia Siro, Mohammad Aliannejadi, Nick Craswell, Charles L. A. Clarke, Guglielmo Faggioli, Bhaskar Mitra, Paul Thomas, and Emine Yilmaz. 2024 · 2024
Later among the works it cites.
AI models collapse when trained on recursively generated data
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross J. Anderson, and Yarin Gal. 2024 · 2024
Later among the works it cites.
A Survey on Recent Advances in Conversational Data Generation
Heydar Soudani, Roxana Petcu, Evangelos Kanoulas, and Faegheh Hasibi. 2024 · 2024
Later among the works it cites.
Large Language Models are Inconsistent and Biased Evaluators
Rickard Stureborg, Dimitris Alikaniotis, and Yoshi Suhara. 2024 · 2024
Later among the works it cites.
Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP ’24) . 17086–17105
Tu Vu, Kalpesh Krishna, Salaheddin Alzubi, Chris Tar, Manaal Faruqui, and Yun-Hsuan Sung. 2024 · 2024
Later among the works it cites.
Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (ACL ’24) . 15474–15492
Wenda Xu, Guanglei Zhu, Xuandong Zhao, Liangming Pan, Lei Li, and William Wang. 2024 · 2024
Later among the works it cites.
The FACTS Grounding Leaderboard: Benchmarking LLMs’ Ability to Ground Responses to Long-Form Input
Alon Jacovi, Andrew Wang, Chris Alberti, Connie Tao, Jon Lipovetz, Kate Olszewska, Lukas Haas, Michelle Liu, Nate Keating, Adam Bloniarz, Carl Saroufim, Corey Fry, Dror Marcus, Doron Kukliansky, Gaurav Singh Tomar, James Swirhun, Jinwei Xing, Lily Wang, Madhu Gurumurthy, Michael Aaron, Moran Ambar, Rachana Fellinger, Rui Wang, Zizhao Zhang, Sasha Goldshtein, and Dipanjan Das. 2025 · 2025
Closest in time.
The Widespread Adoption of Large Language Model-Assisted Writing Across Society
Weixin Liang, Yaohui Zhang, Mihai Codreanu, Jiayu Wang, Hancheng Cao, and James Zou. 2025 · 2025
Closest in time.
SynDL: A Large-Scale Synthetic Test Collection for Passage Retrieval
Hossein A. Rahmani, Xi Wang, Emine Yilmaz, Nick Craswell, Bhaskar Mitra, and Paul Thomas. 2025 · 2025
Closest in time.
Automated Query-Product Relevance Labeling using Large Language Models for E-commerce Search
Jayant Sachdev, Sean D. Rosario, Abhijeet Phatak, He Wen, Swati Kirti, and Chittaranjan Tripathy. 2025 · 2025
Closest in time.
Don’t Use LLMs to Make Relevance Judgments
Ian Soboroff. 2025 · 2025
Closest in time.
Manveer Singh Tamber and Jimmy Lin. 2025 · 2025
Closest in time.