Fetching the paper…
Reading the bibliography…
We introduce FreshStack, a holistic framework for automatically building information retrieval (IR) evaluation benchmarks by incorporating challenging questions and answers.
Neural Code Search Evaluation Dataset
Hongyu Li, Seohyun Kim, and Satish Chandra. 2019 · 1908
Earlier work this paper cites.
CodeSearchNet Challenge: Evaluating the State of Semantic Code Search
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. 2019 · 1909
Earlier work this paper cites.
On the Evaluation of Machine-Generated Reports
James Mayfield, Eugene Yang, Dawn J. Lawrie, Sean MacAvaney, Paul McNamee, Douglas W. Oard, Luca Soldaini, Ian Soboroff, Orion Weller, Efsun Selin Kayi, Kate Sanders, Marc Mason, and Noah Hibbler. 2024 · 1915
Earlier work this paper cites.
Large Language Models can Accurately Predict Searcher Preferences
Paul Thomas, Seth Spielman, Nick Craswell, and Bhaskar Mitra. 2024 · 1940
Earlier work this paper cites.
A Workbench for Autograding Retrieve/Generate Systems
Laura Dietz. 2024 · 1972
Earlier work this paper cites.
The use of MMR, diversity-based reranking for reordering documents and producing summaries
Jaime Carbonell and Jade Goldstein. 1998 · 1998
Earlier work this paper cites.
Overview of the TREC 2003 Question Answering Track
Ellen M. Voorhees. 2003 · 2003
Earlier work this paper cites.
Evaluating Content Selection in Summarization: The Pyramid Method
Ani Nenkova and Rebecca Passonneau. 2004 · 2004
Earlier work this paper cites.
Automatically Evaluating Answers to Definition Questions
Jimmy Lin and Dina Demner-Fushman. 2005 · 2005
Earlier work this paper cites.
Methods for automatically evaluating answers to complex questions
Jimmy Lin and Dina Demner-Fushman. 2006 · 2006
Earlier work this paper cites.
Novelty and diversity in information retrieval evaluation
Charles L.A. Clarke, Maheedhar Kolla, Gordon V. Cormack, Olga Vechtomova, Azin Ashkan, Stefan Büttcher, and Ian MacKinnon. 2008 · 2008
Earlier work this paper cites.
I Come Not To Bury Cranfield, but to Praise It
Ellen Voorhees. 2009 · 2009
Earlier work this paper cites.
IR system evaluation using nugget-based test collections
Virgil Pavlu, Shahzad Rajput, Peter B. Golbus, and Javed A. Aslam. 2012 · 2012
Earlier work this paper cites.
CQADupStack: A Benchmark Data Set for Community Question-Answering Research
Doris Hoogeveen, Karin M. Verspoor, and Timothy Baldwin. 2015 · 2015
Earlier work this paper cites.
Search Result Diversification
Rodrygo L. T. Santos, Craig Macdonald, and Iadh Ounis. 2015 · 2015
Earlier work this paper cites.
The Influence of Topic Difficulty, Relevance Level, and Document Ordering on Relevance Judging
Tadele T. Damessie, Falk Scholer, and J. Shane Culpepper. 2016 · 2016
Earlier work this paper cites.
MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016 · 2016
Earlier work this paper cites.
TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018 · 2018
Earlier work this paper cites.
Natural Questions: a Benchmark for Question Answering Research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur P. Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019 · 2019
Earlier work this paper cites.
Retrieval Augmented Language Model Pre-Training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020 · 2020
Earlier work this paper cites.
Generalization through Memorization: Nearest Neighbor Language Models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020 · 2020
Earlier work this paper cites.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020 · 2020
Earlier work this paper cites.
XOR QA: Cross-lingual Open-Retrieval Question Answering
Akari Asai, Jungo Kasai, Jonathan Clark, Kenton Lee, Eunsol Choi, and Hannaneh Hajishirzi. 2021 · 2021
Earlier work this paper cites.
Overview of the TREC 2021 deep learning track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Jimmy Lin. 2022 · 2021
Earlier work this paper cites.
Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering
Gautier Izacard and Edouard Grave. 2021 · 2021
Earlier work this paper cites.
Pyserini: A Python Toolkit for Reproducible Information Retrieval Research with Sparse and Dense Representations
Jimmy Lin, Xueguang Ma, Sheng-Chieh Lin, Jheng-Hong Yang, Ronak Pradeep, and Rodrigo Frassetto Nogueira. 2021 · 2021
Cited alongside, same era.
CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin Clement, Dawn Drain, Daxin Jiang, Duyu Tang, Ge Li, Lidong Zhou, Linjun Shou, Long Zhou, Michele Tufano, MING GONG, Ming Zhou, Nan Duan, Neel Sundaresan, Shao Kun Deng, Shengyu Fu, and Shujie LIU. 2021 · 2021
Cited alongside, same era.
BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021 · 2021
Cited alongside, same era.
Improving Language Models by Retrieving from Trillions of Tokens
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Chris Jones, Albin Cassirer, Andy Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack W. Rae, Erich Elsen, and Laurent Sifre. 2022 · 2022
Cited alongside, same era.
AttributionBench: How Hard is Automatic Attribution Evaluation?
Yifei Li, Xiang Yue, Zeyi Liao, and Huan Sun. 2024b · 2024
Later among the works it cites.
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models
Cheng Niu, Yuanhao Wu, Juno Zhu, Siliang Xu, KaShun Shum, Randy Zhong, Juntong Song, and Tong Zhang. 2024 · 2024
Later among the works it cites.
Hello GPT-4o
OpenAI. 2024 · 2024
Later among the works it cites.
Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework
Ronak Pradeep, Nandan Thakur, Shivani Upadhyay, Daniel Campos, Nick Craswell, and Jimmy Lin. 2024 · 2024
Later among the works it cites.
LLMJudge: LLMs for Relevance Judgments
Hossein A. Rahmani, Emine Yilmaz, Nick Craswell, Bhaskar Mitra, Paul Thomas, Charles L. A. Clarke, Mohammad Aliannejadi, Clemencia Siro, and Guglielmo Faggioli. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Overview of the TREC 2022 deep learning track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, Jimmy Lin, Ellen M. Voorhees, and Ian Soboroff. 2023 · 2022
Cited alongside, same era.
ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. 2022 · 2022
Cited alongside, same era.
Chain-of-Thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022 · 2022
Cited alongside, same era.
Stochastic Retrieval-Conditioned Reranking
Hamed Zamani, Michael Bendersky, Donald Metzler, Honglei Zhuang, and Xuanhui Wang. 2022 · 2022
Cited alongside, same era.
Overview of the TREC 2023 Deep Learning Track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Hossein A. Rahmani, Daniel Campos, Jimmy Lin, Ellen M. Voorhees, and Ian Soboroff. 2024 · 2023
Cited alongside, same era.
Perspectives on Large Language Models for Relevance Judgment
Guglielmo Faggioli, Laura Dietz, Charles L. A. Clarke, Gianluca Demartini, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Benno Stein, and Henning Wachsmuth. 2023 · 2023
Cited alongside, same era.
Precise Zero-Shot Dense Retrieval without Relevance Labels
Luyu Gao, Xueguang Ma, Jimmy Lin, and Jamie Callan. 2023a · 2023
Cited alongside, same era.
Enabling Large Language Models to Generate Text with Citations
Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen. 2023b · 2023
Cited alongside, same era.
Question-Based Retrieval using Atomic Units for Enterprise RAG
Vatsal Raina and Mark Gales. 2024 · 2024
Later among the works it cites.
Researchy Questions: A Dataset of Multi-Perspective, Decompositional Questions for LLM Web Agents
Corby Rosset, Ho-Lam Chung, Guanghui Qin, Ethan C. Chau, Zhuo Feng, Ahmed Awadallah, Jennifer Neville, and Nikhil Rao. 2024 · 2024
Later among the works it cites.
Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models
Yijia Shao, Yucheng Jiang, Theodore Kanell, Peter Xu, Omar Khattab, and Monica Lam. 2024 · 2024
Later among the works it cites.
Observations on Building RAG Systems for Technical Documents
Sumit Soman and Sujoy Roychowdhury. 2024 · 2024
Later among the works it cites.
CoRNStack: High-Quality Contrastive Data for Better Code Ranking
Tarun Suresh, Revanth Gangi Reddy, Yifei Xu, Zach Nussbaum, Andriy Mulyar, Brandon Duderstadt, and Heng Ji. 2024 · 2024
Later among the works it cites.
Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges
Nandan Thakur, Ronak Pradeep, Shivani Upadhyay, Daniel Campos, Nick Craswell, and Jimmy Lin. 2025 · 2024
Later among the works it cites.
FreshLLMs: Refreshing Large Language Models with Search Engine Augmentation
Tu Vu, Mohit Iyyer, Xuezhi Wang, Noah Constant, Jerry Wei, Jason Wei, Chris Tar, Yun-Hsuan Sung, Denny Zhou, Quoc Le, and Thang Luong. 2024 · 2024
Later among the works it cites.
Improving Text Embeddings with Large Language Models
Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2024a · 2024
Later among the works it cites.
Result Diversification in Search and Recommendation: A Survey
Haolun Wu, Yansen Zhang, Chen Ma, Fuyuan Lyu, Bowei He, Bhaskar Mitra, and Xue Liu. 2024 · 2024
Later among the works it cites.
RAR-b: Reasoning as Retrieval Benchmark
Chenghao Xiao, G. Thomas Hudson, and Noura Al Moubayed. 2024 · 2024
Later among the works it cites.
CRAG - Comprehensive RAG Benchmark
Xiao Yang, Kai Sun, Hao Xin, Yushi Sun, Nikita Bhalla, Xiangsen Chen, Sajal Choudhary, Rongze Gui, Ziran Jiang, Ziyu JIANG, Lingkun Kong, Brian Moran, Jiaqi Wang, Yifan Ethan Xu, An Yan, Chenyu Yang, Eting Yuan, Hanwen Zha, Nan Tang, Lei Chen, Nicolas SCHEFFER, Yue Liu, Nirav Shah, Rakesh Wanga, Anuj Kumar, Wen tau Yih, and Xin Luna Dong. 2024 · 2024
Later among the works it cites.
The Shift from Models to Compound AI Systems
Matei Zaharia, Omar Khattab, Lingjiao Chen, Jared Quincy Davis, Heather Miller, Chris Potts, James Zou, Michael Carbin, Jonathan Frankle, Naveen Rao, and Ali Ghodsi. 2024 · 2024
Later among the works it cites.
Jingyi Chen, Songqiang Chen, Jialun Cao, Jiasi Shen, and Shing-Chi Cheung. 2025 · 2025
Closest in time.
Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation
Satyapriya Krishna, Kalpesh Krishna, Anhad Mohananey, Steven Schwarcz, Adam Stambler, Shyam Upadhyay, and Manaal Faruqui. 2025 · 2025
Closest in time.
Ragnarök: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented Generation Track
Ronak Pradeep, Nandan Thakur, Sahel Sharifymoghaddam, Eric Zhang, Ryan Nguyen, Daniel Campos, Nick Craswell, and Jimmy Lin. 2025a · 2025
Closest in time.
Customized Retrieval Augmented Generation and Benchmarking for EDA Tool Documentation QA
Yuan Pu, Zhuolun He, Tairu Qiu, Haoyuan Wu, and Bei Yu. 2025 · 2025
Closest in time.
CLAPnq: Cohesive Long-form Answers from Passages in Natural Questions for RAG systems
Sara Rosenthal, Avirup Sil, Radu Florian, and Salim Roukos. 2025 · 2025
Closest in time.
Shivalika Singh, Yiyang Nan, Alex Wang, Daniel D’souza, Sayash Kapoor, Ahmet Üstün, Sanmi Koyejo, Yuntian Deng, Shayne Longpre, Noah Smith, Beyza Ermis, Marzieh Fadaee, and Sara Hooker. 2025 · 2025
Closest in time.
BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval
Hongjin SU, Howard Yen, Mengzhou Xia, Weijia Shi, Niklas Muennighoff, Han yu Wang, Liu Haisu, Quan Shi, Zachary S Siegel, Michael Tang, Ruoxi Sun, Jinsung Yoon, Sercan O Arik, Danqi Chen, and Tao Yu. 2025 · 2025
Closest in time.
LiveBench: A Challenging, Contamination-Limited LLM Benchmark
Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Benjamin Feuer, Siddhartha Jain, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Sreemanti Dey, Shubh-Agrawal, Sandeep Singh Sandha, Siddartha Venkat Naidu, Chinmay Hegde, Yann LeCun, Tom Goldstein, Willie Neiswanger, and Micah Goldblum. 2025 · 2025
Closest in time.