Fetching the paper…
Reading the bibliography…
This paper introduces xRAG, an innovative context compression method tailored for retrieval-augmented generation.
Semantic parsing on freebase from question-answer pairs
Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang · 2013
Earlier work this paper cites.
WikiQA: A challenge dataset for open-domain question answering
Yi Yang, Wen-tau Yih, and Christopher Meek · 2015
Earlier work this paper cites.
Reading wikipedia to answer open-domain questions, 2017
Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes · 2017
Earlier work this paper cites.
Imitation learning: A survey of learning methods
Ahmed Hussein, Mohamed Medhat Gaber, Eyad Elyan, and Chrisina Jayne · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension, 2017
Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
The narrativeqa reading comprehension challenge, 2017
Tomáš Kočiský, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, and Edward Grefenstette · 2017
Earlier work this paper cites.
Get to the point: Summarization with pointer-generator networks, 2017
Abigail See, Peter J. Liu, and Christopher D. Manning · 2017
Earlier work this paper cites.
Ms marco: A human generated machine reading comprehension dataset, 2018
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosenberg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang · 2018
Earlier work this paper cites.
Self-imitation learning
Junhyuk Oh, Yijie Guo, Satinder Singh, and Honglak Lee · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for squad, 2018
Pranav Rajpurkar, Robin Jia, and Percy Liang · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering, 2018
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning · 2018
Earlier work this paper cites.
Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs, 2019
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner · 2019
Earlier work this paper cites.
Samsum corpus: A human-annotated dialogue dataset for abstractive summarization
Bogdan Gliwa, Iwona Mochol, Maciej Biesek, and Aleksander Wawer · 2019
Earlier work this paper cites.
FreebaseQA: A new factoid QA data set matching trivia-style question-answer pairs with Freebase
Kelvin Jiang, Dekun Wu, and Hui Jiang · 2019
Earlier work this paper cites.
PubMedQA: A dataset for biomedical research question answering
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu · 2019
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov · 2019
Earlier work this paper cites.
Coqa: A conversational question answering challenge, 2019
Siva Reddy, Danqi Chen, and Christopher D. Manning · 2019
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge, 2019
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2019
Earlier work this paper cites.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2020
Earlier work this paper cites.
Realm: Retrieval-augmented language model pre-training, 2020
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang · 2020
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering, 2020
Vladimir Karpukhin, Barlas Oğuz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen tau Yih · 2020
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models, 2020
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis · 2020
Earlier work this paper cites.
Getting closer to ai complete question answering: A set of prerequisite real tasks
Anna Rogers, Olga Kovaleva, Matthew Downey, and Anna Rumshisky · 2020
Earlier work this paper cites.
Dialogsum: A real-life scenario dialogue summarization dataset, 2021
Yulong Chen, Yang Liu, Liang Chen, and Yue Zhang · 2021
Earlier work this paper cites.
Leveraging passage retrieval with generative models for open domain question answering, 2021
Gautier Izacard and Edouard Grave · 2021
Earlier work this paper cites.
Nearest neighbor machine translation, 2021
Urvashi Khandelwal, Angela Fan, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis · 2021
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks, 2021
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela · 2021
Earlier work this paper cites.
Kilt: a benchmark for knowledge intensive language tasks, 2021
Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, Vassilis Plachouras, Tim Rocktäschel, and Sebastian Riedel · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Earlier work this paper cites.
Improving language models by retrieving from trillions of tokens, 2022
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Chris Jones, Albin Cassirer, Andy Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack W. Rae, Erich Elsen, and Laurent Sifre · 2022
Cited alongside, same era.
Neural machine translation with contrastive translation memories, 2022
Xin Cheng, Shen Gao, Lemao Liu, Dongyan Zhao, and Rui Yan · 2022
Cited alongside, same era.
Atlas: Few-shot learning with retrieval augmented language models, 2022
Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave · 2022
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods, 2022
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Cited alongside, same era.
Colbertv2: Effective and efficient retrieval via lightweight late interaction, 2022
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia · 2022
Nonparametric masked language modeling, 2023
Sewon Min, Weijia Shi, Mike Lewis, Xilun Chen, Wen tau Yih, Hannaneh Hajishirzi, and Luke Zettlemoyer · 2023
Later among the works it cites.
Text embeddings reveal (almost) as much as text, 2023
John X. Morris, Volodymyr Kuleshov, Vitaly Shmatikov, and Alexander M. Rush · 2023
Later among the works it cites.
Mteb: Massive text embedding benchmark, 2023
Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers · 2023
Later among the works it cites.
Replug: Retrieval-augmented black-box language models, 2023
Weijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Rich James, Mike Lewis, Luke Zettlemoyer, and Wen tau Yih · 2023
Later among the works it cites.
How far can camels go? exploring the state of instruction tuning on open resources, 2023
Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, David Wadden, Kelsey MacMillan, Noah A. Smith, Iz Beltagy, and Hannaneh Hajishirzi · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning by distilling context, 2022
Charlie Snell, Dan Klein, and Ruiqi Zhong · 2022
Cited alongside, same era.
Training data is more valuable than you think: A simple and effective method by retrieving from training data, 2022
Shuohang Wang, Yichong Xu, Yuwei Fang, Yang Liu, Siqi Sun, Ruochen Xu, Chenguang Zhu, and Michael Zeng · 2022
Cited alongside, same era.
Finetuned language models are zero-shot learners, 2022
Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le · 2022
Cited alongside, same era.
Training language models with memory augmentation, 2022
Zexuan Zhong, Tao Lei, and Danqi Chen · 2022
Cited alongside, same era.
Retrieval-based language models and applications
Akari Asai, Sewon Min, Zexuan Zhong, and Danqi Chen · 2023
Cited alongside, same era.
Self-rag: Learning to retrieve, generate, and critique through self-reflection, 2023
Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi · 2023
Cited alongside, same era.
Decouple knowledge from parameters for plug-and-play language modeling, 2023
Xin Cheng, Yankai Lin, Xiuying Chen, Dongyan Zhao, and Rui Yan · 2023
Cited alongside, same era.
Zhiruo Wang, Jun Araki, Zhengbao Jiang, Md Rizwan Parvez, and Graham Neubig · 2023
Later among the works it cites.
C-pack: Packaged resources to advance general chinese embedding, 2023
Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighoff · 2023
Later among the works it cites.
Recomp: Improving retrieval-augmented lms with compression and selective augmentation, 2023
Fangyuan Xu, Weijia Shi, and Eunsol Choi · 2023
Later among the works it cites.
Making retrieval-augmented language models robust to irrelevant context, 2023
Ori Yoran, Tomer Wolfson, Ori Ram, and Jonathan Berant · 2023
Later among the works it cites.
Generate rather than retrieve: Large language models are strong context generators, 2023
Wenhao Yu, Dan Iter, Shuohang Wang, Yichong Xu, Mingxuan Ju, Soumya Sanyal, Chenguang Zhu, Michael Zeng, and Meng Jiang · 2023
Later among the works it cites.
Focus on the core: Efficient attention via pruned token compression for document classification
Jungmin Yun, Mihyeon Kim, and Youngbin Kim · 2023
Later among the works it cites.
Reliable, adaptable, and attributable language models with retrieval, 2024
Akari Asai, Zexuan Zhong, Danqi Chen, Pang Wei Koh, Luke Zettlemoyer, Hannaneh Hajishirzi, and Wen tau Yih · 2024
Closest in time.
Quebec automobile insurance question-answering with retrieval-augmented generation
David Beauchemin, Zachary Gagnon, and Ricahrd Khoury · 2024
Closest in time.
Retrieval-augmented generation for large language models: A survey, 2024
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, and Haofen Wang · 2024
Closest in time.
Mixtral of experts, 2024
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2024
Closest in time.
Prismatic vlms: Investigating the design space of visually-conditioned language models, 2024
Siddharth Karamcheti, Suraj Nair, Ashwin Balakrishna, Percy Liang, Thomas Kollar, and Dorsa Sadigh · 2024
Closest in time.
Synthetic data (almost) from scratch: Generalized instruction tuning for language models, 2024
Haoran Li, Qingxiu Dong, Zhengyang Tang, Chaojun Wang, Xingxing Zhang, Haoyang Huang, Shaohan Huang, Xiaolong Huang, Zeqiang Huang, Dongdong Zhang, Yuxian Gu, Xin Cheng, Xun Wang, Si-Qing Chen, Li Dong, Wei Lu, Zhifang Sui, Benyou Wang, Wai Lam, and Furu Wei · 2024
Closest in time.
Prompt compression for large language models: A survey, 2024
Zongqian Li, Yinhong Liu, Yixuan Su, and Nigel Collier · 2024
Closest in time.
500xcompressor: Generalized prompt compression for large language models, 2024
Zongqian Li, Yixuan Su, and Nigel Collier · 2024
Closest in time.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2024
Closest in time.
Leveraging passage embeddings for efficient listwise reranking with large language models, 2024
Qi Liu, Bo Wang, Nan Wang, and Jiaxin Mao · 2024
Closest in time.
Deepseek-vl: Towards real-world vision-language understanding, 2024
Haoyu Lu, Wen Liu, Bo Zhang, Bingxuan Wang, Kai Dong, Bo Liu, Jingxiang Sun, Tongzheng Ren, Zhuoshu Li, Hao Yang, Yaofeng Sun, Chengqi Deng, Hanwei Xu, Zhenda Xie, and Chong Ruan · 2024
Closest in time.
An empirical study of catastrophic forgetting in large language models during continual fine-tuning, 2024
Yun Luo, Zhen Yang, Fandong Meng, Yafu Li, Jie Zhou, and Yue Zhang · 2024
Closest in time.
Learning to compress prompts with gist tokens, 2024
Jesse Mu, Xiang Lisa Li, and Noah Goodman · 2024
Closest in time.
Sfr-rag: Towards contextually faithful llms
Xuan-Phi Nguyen, Shrey Pandit, Senthil Purushwalkam, Austin Xu, Hailin Chen, Yifei Ming, Zixuan Ke, Silvio Savarese, Caiming Xong, and Shafiq Joty · 2024
Closest in time.
Llmlingua-2: Data distillation for efficient and faithful task-agnostic prompt compression, 2024
Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Menglin Xia, Xufang Luo, Jue Zhang, Qingwei Lin, Victor Rühle, Yuqing Yang, Chin-Yew Lin, H. Vicky Zhao, Lili Qiu, and Dongmei Zhang · 2024
Closest in time.
Text embeddings by weakly-supervised contrastive pre-training, 2024
Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei · 2024
Closest in time.
Compact: Compressing retrieved documents actively for question answering, 2024
Chanwoong Yoon, Taewhoo Lee, Hyeon Hwang, Minbyul Jeong, and Jaewoo Kang · 2024
Closest in time.
Funnelrag: A coarse-to-fine progressive retrieval paradigm for rag
Xinping Zhao, Yan Zhong, Zetian Sun, Xinshuo Hu, Zhenyu Liu, Dongfang Li, Baotian Hu, and Min Zhang · 2024
Closest in time.