Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are increasingly integrated into real-world personalized applications through retrieval-augmented generation (RAG) mechanisms to supplement their responses with domain-specific knowledge.
Nearest neighbor pattern classification
Thomas Cover and Peter Hart · 1967
Earlier work this paper cites.
Introduction to mathematical statistics
Leopold Schmetterer · 2012
Earlier work this paper cites.
Ms marco: A human generated machine reading comprehension dataset
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, et al · 2016
Earlier work this paper cites.
Membership inference attacks against machine learning models
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov · 2017
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning · 2018
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al · 2019
Earlier work this paper cites.
Rodrigo Nogueira and Kyunghyun Cho · 2019
Earlier work this paper cites.
Poly-encoders: Transformer architectures and pre-training strategies for fast and accurate multi-sentence scoring
Samuel Humeau, Kurt Shuster, Marie-Anne Lachaux, and Jason Weston · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
Relevance-guided supervision for openqa with colbert
Omar Khattab, Christopher Potts, and Matei Zaharia · 2021
Earlier work this paper cites.
Approximate nearest neighbor negative contrastive learning for dense text retrieval
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N Bennett, Junaid Ahmed, and Arnold Overwijk · 2021
Earlier work this paper cites.
Unsupervised dense information retrieval with contrastive learning
Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave · 2022
Earlier work this paper cites.
Untargeted backdoor watermark: Towards harmless and stealthy dataset copyright protection
Yiming Li, Yang Bai, Yong Jiang, Yong Yang, Shu-Tao Xia, and Bo Li · 2022
Earlier work this paper cites.
Defending against model stealing via verifying embedded external features
Yiming Li, Linghui Zhu, Xiaojun Jia, Yong Jiang, Shu-Tao Xia, and Xiaochun Cao · 2022
Earlier work this paper cites.
Mteb: Massive text embedding benchmark
Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Detecting language model attacks with perplexity
Gabriel Alon and Michael Kamfonas · 2023
Earlier work this paper cites.
Towards next-generation intelligent assistants leveraging llm techniques
Xin Luna Dong, Seungwhan Moon, Yifan Ethan Xu, Kshitiz Malik, and Zhou Yu · 2023
Earlier work this paper cites.
Chatgpt for (finance) research: The bananarama conjecture
Michael Dowling and Brian Lucey · 2023
Earlier work this paper cites.
Jina embeddings 2: 8192-token general-purpose text embeddings for long documents
Michael Günther, Jackmin Ong, Isabelle Mohr, Alaeddine Abdessalem, Tanguy Abel, Mohammad Kalim Akram, Susana Guzman, Georgios Mastrapas, Saba Sturua, Bo Wang, et al · 2023
Earlier work this paper cites.
Domain watermark: Effective and harmless dataset copyright protection is closed at hand
Junfeng Guo, Yiming Li, Lixu Wang, Shu-Tao Xia, Heng Huang, Cong Liu, and Bo Li · 2023
Cited alongside, same era.
Loggpt: Log anomaly detection via gpt, 2023
Xiao Han, Shuhan Yuan, and Mohamed Trabelsi · 2023
Cited alongside, same era.
Certifying llm safety against adversarial prompting
Aounon Kumar, Chirag Agarwal, Suraj Srinivas, Aaron Jiaxun Li, Soheil Feizi, and Himabindu Lakkaraju · 2023
Cited alongside, same era.
Black-box dataset ownership verification via backdoor watermarking
Yiming Li, Mingyan Zhu, Xue Yang, Yong Jiang, Tao Wei, and Shu-Tao Xia · 2023
Cited alongside, same era.
Did you train on my dataset? towards public dataset protection with clean-label backdoor watermarking
Ruixiang Tang, Qizhang Feng, Ninghao Liu, Fan Yang, and Xia Hu · 2023
Cited alongside, same era.
Data extraction attacks in retrieval-augmented generation via backdoors
Yuefeng Peng, Junda Wang, Hong Yu, and Amir Houmansadr · 2024
Later among the works it cites.
Large language models for forecasting and anomaly detection: A systematic literature review, 2024
Jing Su, Chufeng Jiang, Xin Jin, Yuxin Qiao, Tingsong Xiao, Hongda Ma, Rong Wei, Zhi Jing, Jiajun Xu, and Junhong Lin · 2024
Later among the works it cites.
Redpajama: an open dataset for training large language models
Maurice Weber, Daniel Y. Fu, Quentin Anthony, Yonatan Oren, Shane Adams, Anton Alexandrov, Xiaozhong Lyu, Huu Nguyen, Xiaozhe Yao, Virginia Adams, Ben Athiwaratkun, Rahul Chalamala, Kezhen Chen, Max Ryabinin, Tri Dao, Percy Liang, Christopher Ré, Irina Rish, and Ce Zhang · 2024
Later among the works it cites.
Pointncbw: Towards dataset ownership verification for point clouds via negative clean-label backdoor watermark
Cheng Wei, Yang Wang, Kuofeng Gao, Shuo Shao, Yiming Li, Zhibo Wang, and Zhan Qin · 2024
Later among the works it cites.
Retrieval-augmented generation for natural language processing: A survey
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Cited alongside, same era.
Watermarking graph neural networks based on backdoor attacks
Jing Xu, Stefanos Koffas, Oğuzhan Ersoy, and Stjepan Picek · 2023
Cited alongside, same era.
Enhancing financial sentiment analysis via retrieval augmented large language models
Boyu Zhang, Hongyang Yang, Tianyu Zhou, Muhammad Ali Babar, and Xiao-Yang Liu · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Cited alongside, same era.
Agentpoison: Red-teaming llm agents via poisoning memory or knowledge bases
Zhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song, and Bo Li · 2024
Cited alongside, same era.
Trojanrag: Retrieval-augmented generation can be backdoor driver in large language models
Pengzhou Cheng, Yidong Ding, Tianjie Ju, Zongru Wu, Wei Du, Ping Yi, Zhuosheng Zhang, and Gongshen Liu · 2024
Cited alongside, same era.
Stav Cohen, Ron Bitton, and Ben Nassi · 2024
Cited alongside, same era.
Shangyu Wu, Ying Xiong, Yufei Cui, Haolun Wu, Can Chen, Ye Yuan, Lianming Huang, Xue Liu, Tei-Wei Kuo, Nan Guan, et al · 2024
Later among the works it cites.
Badchain: Backdoor chain-of-thought prompting for large language models
Zhen Xiang, Fengqing Jiang, Zidi Xiong, Bhaskar Ramasubramanian, Radha Poovendran, and Bo Li · 2024
Later among the works it cites.
Promptcare: Prompt copyright protection by watermark injection and verification
Hongwei Yao, Jian Lou, Zhan Qin, and Kui Ren · 2024
Later among the works it cites.
Almanac—retrieval-augmented language models for clinical medicine
Cyril Zakka, Rohan Shad, Akash Chaurasia, Alex R Dalal, Jennifer L Kim, Michael Moor, Robyn Fong, Curran Phillips, Kevin Alexander, Euan Ashley, et al · 2024
Later among the works it cites.
The good and the bad: Exploring privacy issues in retrieval-augmented generation (RAG)
Shenglai Zeng, Jiankun Zhang, Pengfei He, Yiding Liu, Yue Xing, Han Xu, Jie Ren, Yi Chang, Shuaiqiang Wang, Dawei Yin, and Jiliang Tang · 2024
Later among the works it cites.
Mitigating the privacy issues in retrieval-augmented generation (rag) via pure synthetic data
Shenglai Zeng, Jiankun Zhang, Pengfei He, Jie Ren, Tianqi Zheng, Hanqing Lu, Han Xu, Hui Liu, Yue Xing, and Jiliang Tang · 2024
Later among the works it cites.
Sok: Dataset copyright auditing in machine learning systems
Linkang Du, Xuanru Zhou, Min Chen, Chusong Zhang, Zhou Su, Peng Cheng, Jiming Chen, and Zhikun Zhang · 2025
Closest in time.
When backdoors speak: Understanding llm backdoor attacks through model-generated explanations
Huaizhi Ge, Yiming Li, Qifan Wang, Yongfeng Zhang, and Ruixiang Tang · 2025
Closest in time.
Towards label-only membership inference attack against pre-trained large language models
Yu He, Boheng Li, Liu Liu, Zhongjie Ba, Wei Dong, Yiming Li, Zhan Qin, Kui Ren, and Chun Chen · 2025
Closest in time.
Ward: Provable RAG dataset inference via LLM watermarks
Nikola Jovanović, Robin Staab, Maximilian Baader, and Martin Vechev · 2025
Closest in time.
Towards reliable verification of unauthorized data usage in personalized text-to-image diffusion models
Boheng Li, Yanhao Wei, Yankai Fu, Zhenting Wang, Yiming Li, Jie Zhang, Run Wang, and Tianwei Zhang · 2025
Closest in time.
Move: Effective and harmless ownership verification via embedded external features
Yiming Li, Linghui Zhu, Xiaojun Jia, Yang Bai, Yong Jiang, Shu-Tao Xia, and Xiaochun Cao · 2025
Closest in time.
Follow my instruction and spill the beans: Scalable data extraction from retrieval-augmented generation systems
Zhenting Qi, Hanlin Zhang, Eric P. Xing, Sham M. Kakade, and Himabindu Lakkaraju · 2025
Closest in time.
Explanation as a watermark: Towards harmless and multi-bit model ownership verification via watermarking feature attribution
Shuo Shao, Yiming Li, Hongwei Yao, Yiling He, Zhan Qin, and Kui Ren · 2025
Closest in time.
Probe before you talk: Towards black-box defense against backdoor unalignment for large language models
Biao Yi, Tiansheng Huang, Sishuo Chen, Tong Li, Zheli Liu, Chu Zhixuan, and Yiming Li · 2025
Closest in time.
Poisonedrag: Knowledge poisoning attacks to retrieval-augmented generation of large language models
Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia · 2025
Closest in time.