Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are constrained by outdated information and a tendency to generate incorrect data, commonly referred to as "hallucinations." Retrieval-Augmented Generation (RAG) addresses these limitations by combining the strengths of retrieval-based methods and generative models.
Some methods for classification and analysis of multivariate observations
James MacQueen et al · 1967
Earlier work this paper cites.
Interpolated estimation of markov source parameters from sparse data
Frederick Jelinek · 1980
Earlier work this paper cites.
Ms marco: A human generated machine reading comprehension dataset
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, et al · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih · 2020
Earlier work this paper cites.
Approximate nearest neighbor negative contrastive learning for dense text retrieval
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N Bennett, Junaid Ahmed, and Arnold Overwijk · 2020
Earlier work this paper cites.
Beir: A heterogenous benchmark for zero-shot evaluation of information retrieval models
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych · 2021
Earlier work this paper cites.
Trec-covid: constructing a pandemic information retrieval test collection
Ellen Voorhees, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, William R Hersh, Kyle Lo, Kirk Roberts, Ian Soboroff, and Lucy Lu Wang · 2021
Earlier work this paper cites.
Retrieval augmented code generation and summarization
Md Rizwan Parvez, Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang · 2021
Earlier work this paper cites.
Unsupervised dense information retrieval with contrastive learning
Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave · 2021
Earlier work this paper cites.
Wikiasp: A dataset for multi-domain aspect-based summarization
Hiroaki Hayashi, Prashant Budania, Peng Wang, Chris Ackerson, Raj Neervannan, and Graham Neubig · 2021
Earlier work this paper cites.
Trojtext: Test-time invisible textual trojan insertion
Qian Lou, Yepeng Liu, and Bo Feng · 2022
Earlier work this paper cites.
Halueval: A large-scale hallucination evaluation benchmark for large language models
Junyi Li, Xiaoxue Cheng, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen · 2023
Earlier work this paper cites.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung · 2023
Earlier work this paper cites.
Potential for gpt technology to optimize future clinical decision-making using retrieval-augmented generation
Calvin Wang, Joshua Ong, Chara Wang, Hannah Ong, Rebekah Cheng, and Dennis Ong · 2023
Earlier work this paper cites.
Making llms worth every penny: Resource-limited text classification in banking
Lefteris Loukas, Ilias Stogiannidis, Odysseas Diamantopoulos, Prodromos Malakasiotis, and Stavros Vassos · 2023
Earlier work this paper cites.
Chain of reference prompting helps llm to think like a lawyer
Aditya Kuppa, Nikon Rasumov-Rahe, and Marc Voses · 2023
Cited alongside, same era.
Wikichat: Stopping the hallucination of large language model chatbots by few-shot grounding on wikipedia
Sina Semnani, Violet Yao, Heidi Zhang, and Monica Lam · 2023
Cited alongside, same era.
Bluebot–jetblue’s unified llm
YouTube · 2023
Cited alongside, same era.
Poisoning retrieval corpora by injecting adversarial passages
Zexuan Zhong, Ziqing Huang, Alexander Wettig, and Danqi Chen · 2023
Cited alongside, same era.
Enhancing financial sentiment analysis via retrieval augmented large language models
Boyu Zhang, Hongyang Yang, Tianyu Zhou, Muhammad Ali Babar, and Xiao-Yang Liu · 2023
Cited alongside, same era.
Improving the domain adaptation of retrieval augmented generation (rag) models for open domain question answering
Poisonedrag: Knowledge poisoning attacks to retrieval-augmented generation of large language models
Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia · 2024
Closest in time.
Sukmin Cho, Soyeong Jeong, Jeongyeon Seo, Taeho Hwang, and Jong C Park · 2024
Closest in time.
Trojllm: A black-box trojan prompt attack on large language models
Jiaqi Xue, Mengxin Zheng, Ting Hua, Yilin Shen, Yepeng Liu, Ladislau Bölöni, and Qian Lou · 2024
Closest in time.
Test-time backdoor attacks on multimodal large language models
Dong Lu, Tianyu Pang, Chao Du, Qian Liu, Xianjun Yang, and Min Lin · 2024
Closest in time.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shamane Siriwardhana, Rivindu Weerasekera, Elliott Wen, Tharindu Kaluarachchi, Rajib Rana, and Suranga Nanayakkara · 2023
Cited alongside, same era.
Trojbits: A hardware aware inference-time attack on transformer-based language models
Mansour Al Ghanim, Muhammad Santriaji, Qian Lou, and Yan Solihin · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson · 2023
Cited alongside, same era.
Trojfsp: Trojan insertion in few-shot prompt tuning
Mengxin Zheng, Jiaqi Xue, Xun Chen, YanShan Wang, Qian Lou, and Lei Jiang · 2023
Cited alongside, same era.
Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz · 2023
Cited alongside, same era.
Prompt injection attack against llm-integrated applications
Yi Liu, Gelei Deng, Yuekang Li, Kailong Wang, Tianwei Zhang, Yepang Liu, Haoyu Wang, Yan Zheng, and Yang Liu · 2023
Cited alongside, same era.
Backdooring instruction-tuned large language models with virtual prompt injection
Jun Yan, Vikas Yadav, Shiyang Li, Lichang Chen, Zheng Tang, Hai Wang, Vijay Srinivasan, Xiang Ren, and Hongxia Jin · 2023
Cited alongside, same era.
Closest in time.
Backdoor attacks on dense passage retrievers for disseminating misinformation
Quanyu Long, Yue Deng, LeiLei Gan, Wenya Wang, and Sinno Jialin Pan · 2024
Closest in time.
https://python.langchain.com/v0.1/docs/integrations/text_embedding/llamacpp/ , 2024
Llama-cpp text embeddings · 2024
Closest in time.
Multi-task contrastive learning for 8192-token bilingual text embeddings
Isabelle Mohr, Markus Krimmel, Saba Sturua, Mohammad Kalim Akram, Andreas Koukounas, Michael Günther, Georgios Mastrapas, Vinit Ravishankar, Joan Fontanals Martínez, Feng Wang, et al · 2024
Closest in time.
https://huggingface.co/jinaai/jina-embeddings-v2-base-de , 2024
Huggingface: jinaai/jina-embeddings-v2-base-de · 2024
Closest in time.
https://huggingface.co/facebook/contriever , 2024
Huggingface: facebook/contriever · 2024
Closest in time.
https://www.llamaindex.ai/ , 2024
Llamaindex · 2024
Closest in time.
https://www.langchain.com/ , 2024
Langchain · 2024
Closest in time.
https://www.wikipedia.org/ , 2024
Wikipedia · 2024
Closest in time.
https://www.reddit.com/ , 2024
Reddit · 2024
Closest in time.
https://huggingface.co/datasets/wikitext , 2024
Wikidataset: Huggingface · 2024
Closest in time.
https://www.commoncrawl.org/ , 2024
Common crawl · 2024
Closest in time.
Lost in the middle: How language models use long contexts
Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang · 2024
Closest in time.
Defending against indirect prompt injection attacks with spotlighting
Keegan Hines, Gary Lopez, Matthew Hall, Federico Zarfati, Yonatan Zunger, and Emre Kiciman · 2024
Closest in time.
Chat with claude
Anthropic · 2024
Closest in time.