Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have enabled new ways to satisfy information needs.
Report on the Need for and Provision of an ‘ideal’ Information Retrieval Test Collection
Karen Sparck Jones and C.J. Van Rijsbergen. 1975 · 1975
Earlier work this paper cites.
The Text REtrieval Conferences (TRECs). In TIPSTER TEXT PROGRAM PHASE III: Proceedings of a Workshop held at Baltimore, Maryland, October 13-15, 1998 . Association for Computational Linguistics, Baltimore, Maryland, USA, 241–273
Ellen M. Voorhees and Donna Harman. 1998 · 1998
Earlier work this paper cites.
How reliable are the results of large-scale information retrieval experiments?. In Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (Melbourne, Australia) (SIGIR ’98) . Association for Computing Machinery, New York, NY, USA, 307–314
Justin Zobel. 1998 · 1998
Earlier work this paper cites.
The effect of pool depth on system evaluation in TREC
Sabrina Keenan, Alan F. Smeaton, and Gary Keogh. 2001 · 2001
Earlier work this paper cites.
Domain-specific informative and indicative summarization for information retrieval. In SIGIR Worshop on Text Summarization
Judith L Klavans, Min-yen Kan, and Kathleen McKeown. 2001 · 2001
Earlier work this paper cites.
Automatic summarization
Inderjeet Mani. 2001 · 2001
Earlier work this paper cites.
The TREC question answering track
Ellen M Voorhees. 2001 · 2001
Earlier work this paper cites.
Cumulated gain-based evaluation of IR techniques
Kalervo Järvelin and Jaana Kekäläinen. 2002 · 2002
Earlier work this paper cites.
From single to multi-document summarization. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics . 457–464
Chin-Yew Lin and Eduard Hovy. 2002 · 2002
Earlier work this paper cites.
Automatic Evaluation of Summaries Using N-gram Co-occurrence Statistics. In Proceedings of the 2003 Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics . 150–157
Chin-Yew Lin and Eduard Hovy. 2003 · 2003
Earlier work this paper cites.
Sentence level discourse parsing using syntactic and lexical information. In Proceedings of the 2003 Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics . 228–235
Radu Soricut and Daniel Marcu. 2003 · 2003
Earlier work this paper cites.
The Effects of Human Variation in DUC Summarization Evaluation. In Text Summarization Branches Out . Association for Computational Linguistics, Barcelona, Spain, 10–17
Donna Harman and Paul Over. 2004 · 2004
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out . Association for Computational Linguistics, Barcelona, Spain, 74–81
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Evaluating Content Selection in Summarization: The Pyramid Method. In Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics: HLT-NAACL 2004 . Association for Computational Linguistics, Boston, Massachusetts, USA, 145–152
Ani Nenkova and Rebecca Passonneau. 2004 · 2004
Earlier work this paper cites.
Automatic evaluation of text coherence: models and representations. In Proceedings of the 19th International Joint Conference on Artificial Intelligence (Edinburgh, Scotland) (IJCAI’05) . Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 1085–1090
Mirella Lapata and Regina Barzilay. 2005 · 2005
Earlier work this paper cites.
Overview of the TREC 2007 Question Answering Track.. In Trec , Vol. 7. 63
Hoa Trang Dang, Diane Kelly, Jimmy Lin, et al · 2007
Earlier work this paper cites.
A comparison of statistical significance tests for information retrieval evaluation. In Proceedings of the Sixteenth ACM Conference on Conference on Information and Knowledge Management (Lisbon, Portugal) (CIKM ’07) . Association for Computing Machinery, New York, NY, USA, 623–632
Mark D. Smucker, James Allan, and Ben Carterette. 2007 · 2007
Earlier work this paper cites.
Beyond SumBasic: Task-focused summarization with sentence simplification and lexical expansion
Lucy Vanderwende, Hisami Suzuki, Chris Brockett, and Ani Nenkova. 2007 · 2007
Earlier work this paper cites.
An Evaluation Framework for Plagiarism Detection. In Coling 2010: Posters , Chu-Ren Huang and Dan Jurafsky (Eds.). Coling 2010 Organizing Committee, Beijing, China, 997–1005
Martin Potthast, Benno Stein, Alberto Barrón-Cedeño, and Paolo Rosso. 2010 · 2010
Earlier work this paper cites.
Automatically Evaluating Text Coherence Using Discourse Relations. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies , Dekang Lin, Yuji Matsumoto, and Rada Mihalcea (Eds.). Association for Computational Linguistics, Portland, Oregon, USA, 997–1006
Ziheng Lin, Hwee Tou Ng, and Min-Yen Kan. 2011 · 2011
Earlier work this paper cites.
Using Rhetorical Structure Theory and Entity Grids to Automatically Evaluate Local Coherence in Texts. In Computational Processing of the Portuguese Language , Jorge Baptista, Nuno Mamede, Sara Candeias, Ivandré Paraboni, Thiago A. S. Pardo, and Maria das Graças Volpe Nunes (Eds.). Springer International Publishing, Cham, 232–243
Márcio de S. Dias, Valéria D. Feltrim, and Thiago Alexandre Salgueiro Pardo. 2014 · 2014
Earlier work this paper cites.
Incremental update summarization: Adaptive sentence selection based on prevalence and novelty. In Proceedings of the 23rd ACM international conference on conference on information and knowledge management . 301–310
Richard McCreadie, Craig Macdonald, and Iadh Ounis. 2014 · 2014
Earlier work this paper cites.
Re-evaluating Automatic Summarization with BLEU and 192 Shades of ROUGE. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing , Lluís Màrquez, Chris Callison-Burch, and Jian Su (Eds.). Association for Computational Linguistics, Lisbon, Portugal, 128–137
Yvette Graham. 2015 · 2015
Earlier work this paper cites.
TREC 2015 Total Recall Track Overview.. In TREC
Adam Roegiest, Gordon V Cormack, Charles LA Clarke, and Maura R Grossman. 2015 · 2015
Earlier work this paper cites.
Abstractive sentence summarization with attentive recurrent neural networks. In Proceedings of the 2016 Conference of the North American chapter of the Association for Computational Linguistics: Human Language Technologies . 93–98
Sumit Chopra, Michael Auli, and Alexander M Rush. 2016 · 2016
Earlier work this paper cites.
TREC 2016 Total Recall Track Overview.. In TREC
Maura R Grossman, Gordon V Cormack, and Adam Roegiest. 2016 · 2016
Earlier work this paper cites.
Overview of TAC KBP 2016 Event Nugget Track.. In TAC
Teruko Mitamura, Zhengzhong Liu, and Eduard H Hovy. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Active Sampling for Large-scale Information Retrieval Evaluation. In Proceedings of the 2017 ACM on Conference on Information and Knowledge Management (Singapore, Singapore) (CIKM ’17) . Association for Computing Machinery, New York, NY, USA, 49–58
Dan Li and Evangelos Kanoulas. 2017 · 2017
Earlier work this paper cites.
Retrieve, rerank and rewrite: Soft template based neural summarization. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 152–161
Ziqiang Cao, Wenjie Li, Sujian Li, and Furu Wei. 2018 · 2018
Earlier work this paper cites.
Delete, retrieve, generate: a simple approach to sentiment and style transfer
Juncen Li, Robin Jia, He He, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
Cross-document, cross-language event coreference annotation using event hoppers. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018)
Zhiyi Song, Ann Bies, Justin Mott, Xuansong Li, Stephanie Strassel, and Christopher Caruso. 2018 · 2018
Cited alongside, same era.
Neural fuzzy repair: Integrating fuzzy matches into neural machine translation. In 57th Annual Meeting of the Association-for-Computational-Linguistics (ACL) . 1800–1809
Bram Bulte and Arda Tezcan. 2019 · 2019
Cited alongside, same era.
Dynamic Sampling Meets Pooling. In Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval (Paris, France) (SIGIR’19) . Association for Computing Machinery, New York, NY, USA, 1217–1220
Gordon V. Cormack, Haotian Zhang, Nimesh Ghelani, Mustafa Abualsaud, Mark D. Smucker, Maura R. Grossman, Shahin Rahbariasl, and Amira Ghenai. 2019 · 2019
Cited alongside, same era.
ELI5: Long Form Question Answering
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019 · 2019
TRUE: Re-evaluating Factual Consistency Evaluation. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , Marine Carpuat, Marie-Catherine de Marneffe, and Ivan Vladimir Meza Ruiz (Eds.). Association for Computational Linguistics, Seattle, United States, 3905–3920
Or Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansy, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, and Yossi Matias. 2022 · 2022
Later among the works it cites.
A survey on retrieval-augmented text generation
Huayang Li, Yixuan Su, Deng Cai, Yan Wang, and Lemao Liu. 2022 · 2022
Later among the works it cites.
Retrieval augmented visual question answering with outside knowledge
Weizhe Lin and Bill Byrne. 2022 · 2022
Later among the works it cites.
LlamaIndex
Jerry Liu. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Academic Plagiarism Detection: A Systematic Literature Review
Tomáš Foltýnek, Norman Meuschke, and Bela Gipp. 2019 · 2019
Cited alongside, same era.
Automated Pyramid Summarization Evaluation. In Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL) , Mohit Bansal and Aline Villavicencio (Eds.). Association for Computational Linguistics, Hong Kong, China, 404–418
Yanjun Gao, Chen Sun, and Rebecca J. Passonneau. 2019 · 2019
Cited alongside, same era.
Text generation with exemplar-based adaptive decoding
Hao Peng, Ankur P Parikh, Manaal Faruqui, Bhuwan Dhingra, and Dipanjan Das. 2019 · 2019
Cited alongside, same era.
Sentence-BERT: Sentence embeddings using siamese BERT-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Cited alongside, same era.
SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019a · 2019
Cited alongside, same era.
BERTScore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 2019
Cited alongside, same era.
REALM: Retrieval-Augmented Language Model Pre-Training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020 · 2020
Cited alongside, same era.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020 · 2020
Cited alongside, same era.
When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories. In Annual Meeting of the Association for Computational Linguistics
Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Hannaneh Hajishirzi, and Daniel Khashabi. 2022 · 2022
Later among the works it cites.
Retrieval-augmented multimodal language modeling
Michihiro Yasunaga, Armen Aghajanyan, Weijia Shi, Rich James, Jure Leskovec, Percy Liang, Mike Lewis, Luke Zettlemoyer, and Wen-tau Yih. 2022 · 2022
Later among the works it cites.
Visualize Before You Write: Imagination-Guided Open-Ended Text Generation
Wanrong Zhu, An Yan, Yujie Lu, Wenda Xu, Xin Eric Wang, Miguel Eckstein, and William Yang Wang. 2022 · 2022
Later among the works it cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Later among the works it cites.
OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, Jenia Jitsev, Simon Kornblith, Pang Wei Koh, Gabriel Ilharco, Mitchell Wortsman, and Ludwig Schmidt. 2023 · 2023
Later among the works it cites.
MegaWika: Millions of reports and their sources across 50 diverse languages
Samuel Barham, Orion Weller, Michelle Yuan, Kenton Murray, Mahsa Yarmohammadi, Zhengping Jiang, Siddharth Vashishtha, Alexander Martin, Anqi Liu, Aaron Steven White, et al · 2023
Later among the works it cites.
BooookScore: A systematic exploration of book-length summarization in the era of LLMs
Yapei Chang, Kyle Lo, Tanya Goyal, and Mohit Iyyer. 2023 · 2023
Later among the works it cites.
Benchmarking Large Language Models in Retrieval-Augmented Generation
Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2023 · 2023
Later among the works it cites.
Enabling Large Language Models to Generate Text with Citations
Tianyu Gao, Ho-Ching Yen, Jiatong Yu, and Danqi Chen. 2023 · 2023
Later among the works it cites.
Evaluating Generative Ad Hoc Information Retrieval
Lukas Gienapp, Harrisen Scells, Niklas Deckers, Janek Bevendorff, Shuai Wang, Johannes Kiesel, Shahbaz Syed, Maik Frobe, Guide Zucoon, Benno Stein, Matthias Hagen, and Martin Potthast. 2023 · 2023
Later among the works it cites.
Open Domain Multi-document Summarization: A Comprehensive Study of Model Brittleness under Retrieval. In Findings of the Association for Computational Linguistics: EMNLP 2023 . 8177–8199
John Giorgi, Luca Soldaini, Bo Wang, Gary Bader, Kyle Lo, Lucy Lu Wang, and Arman Cohan. 2023 · 2023
Later among the works it cites.
REVEAL: Retrieval-augmented visual-language pre-training with multi-source multimodal knowledge memory. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 23369–23379
Ziniu Hu, Ahmet Iscen, Chen Sun, Zirui Wang, Kai-Wei Chang, Yizhou Sun, Cordelia Schmid, David A Ross, and Alireza Fathi. 2023 · 2023
Later among the works it cites.
Evaluating Verifiability in Generative Search Engines
Nelson F. Liu, Tianyi Zhang, and Percy Liang. 2023b · 2023
Later among the works it cites.
SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore
Sewon Min, Suchin Gururangan, Eric Wallace, Hannaneh Hajishirzi, Noah A. Smith, and Luke Zettlemoyer. 2023a · 2023
Later among the works it cites.
FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023b · 2023
Later among the works it cites.
Efficient text-image semantic search: A multi-modal vision-language approach for fashion retrieval
Gianluca Moro, Stefano Salvatori, and Giacomo Frisoni. 2023 · 2023
Later among the works it cites.
Recognizing textual entailment: A review of resources, approaches, applications, and challenges
I Made Suwija Putra, Daniel Siahaan, and Ahmad Saikhu. 2023 · 2023
Later among the works it cites.
Measuring attribution in natural language generation models
Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm, Lora Aroyo, Michael Collins, Dipanjan Das, Slav Petrov, Gaurav Singh Tomar, Iulia Turc, and David Reitter. 2023 · 2023
Later among the works it cites.
ARES: An automated evaluation framework for retrieval-augmented generation systems
Jon Saad-Falcon, Omar Khattab, Christopher Potts, and Matei Zaharia. 2023 · 2023
Later among the works it cites.
RAGAS: Automated Evaluation of Retrieval Augmented Generation
ES Shahul, Jithin James, Luis Espinosa Anke, and Steven Schockaert. 2023 · 2023
Later among the works it cites.
NoMIRACL: Knowing When You Don’t Know for Robust Multilingual Retrieval-Augmented Generation
Nandan Thakur, Luiz Bonifacio, Xinyu Crystina Zhang, Odunayo Ogundepo, Ehsan Kamalloo, David Alfonso-Hermelo, Xiaoguang Li, Qun Liu, Boxing Chen, Mehdi Rezagholizadeh, and Jimmy J. Lin. 2023 · 2023
Later among the works it cites.
A Critical Evaluation of Evaluations for Long-form Question Answering
Fangyuan Xu, Yixiao Song, Mohit Iyyer, and Eunsol Choi. 2023 · 2023
Later among the works it cites.
Automatic Evaluation of Attribution by Large Language Models. In Conference on Empirical Methods in Natural Language Processing
Xiang Yue, Boshi Wang, Kai Zhang, Ziru Chen, Yu Su, and Huan Sun. 2023 · 2023
Later among the works it cites.
ODSum: New Benchmarks for Open Domain Multi-Document Summarization
Yijie Zhou, Kejian Shi, Wencai Zhang, Yixin Liu, Yilun Zhao, and Arman Cohan. 2023 · 2023
Later among the works it cites.
Improved Evaluation Framework for Complex Plagiarism Detection. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , Iryna Gurevych and Yusuke Miyao (Eds.). Association for Computational Linguistics, Melbourne, Australia, 157–162
Anton Belyy, Marina Dubova, and Dmitry Nekrasov. 2018 · 2026
Closest in time.
Evaluation of Different Plagiarism Detection Methods: A Fuzzy MCDM Perspective
Kamal Mansour Jambi, Imtiaz Hussain Khan, and Muazzam Ahmed Siddiqui. 2022 · 2076
Closest in time.