Fetching the paper…
Reading the bibliography…
In recent years, tremendous success has been witnessed in Retrieval-Augmented Generation (RAG), widely used to enhance Large Language Models (LLMs) in domain-specific, knowledge-intensive, and privacy-sensitive tasks.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Entangled watermarks as a defense against model extraction. In 30th USENIX security symposium (USENIX Security 21) . 1937–1954
Hengrui Jia, Christopher A Choquette-Choo, Varun Chandrasekaran, and Nicolas Papernot. 2021 · 1954
Earlier work this paper cites.
CYC: A large-scale investment in knowledge infrastructure
Douglas B Lenat. 1995 · 1995
Earlier work this paper cites.
Natural language watermarking: Design, analysis, and a proof-of-concept implementation. In Information Hiding: 4th International Workshop, IH 2001 Pittsburgh, PA, USA, April 25–27, 2001 Proceedings 4 . Springer, 185–200
Mikhail J Atallah, Victor Raskin, Michael Crogan, Christian Hempelmann, Florian Kerschbaum, Dina Mohamed, and Sanket Naik. 2001 · 2001
Earlier work this paper cites.
Watermarking relational databases. In VLDB’02: Proceedings of the 28th International Conference on Very Large Databases . Elsevier, 155–166
Rakesh Agrawal and Jerry Kiernan. 2002 · 2002
Earlier work this paper cites.
A system for watermarking relational databases. In Proceedings of the 2003 ACM SIGMOD international conference on Management of data . 674–674
Rakesh Agrawal, Peter J Haas, and Jerry Kiernan. 2003 · 2003
Earlier work this paper cites.
Rights protection for relational data. In Proceedings of the 2003 ACM SIGMOD international conference on Management of data . 98–109
Radu Sion, Mikhail Atallah, and Sunil Prabhakar. 2003 · 2003
Earlier work this paper cites.
Tamper detection and localization for categorical data using fragile watermarks. In Proceedings of the 4th ACM workshop on Digital rights management . 73–82
Yingjiu Li, Huiping Guo, and Sushil Jajodia. 2004 · 2004
Earlier work this paper cites.
Watermarking relational database using image. In Proceedings of 2004 International Conference on Machine Learning and Cybernetics (IEEE Cat. No. 04EX826) , Vol. 3. IEEE, 1739–1744
Zhi-hao Zhang, Xiao-Ming Jin, Jian-Min Wang, and De-Yi Li. 2004 · 2004
Earlier work this paper cites.
The hiding virtues of ambiguity: quantifiably resilient watermarking of natural language text through synonym substitutions. In Proceedings of the 8th workshop on Multimedia and security . 164–174
Umut Topkara, Mercan Topkara, and Mikhail J Atallah. 2006 · 2006
Earlier work this paper cites.
Watermarking relational databases using optimization-based techniques
Mohamed Shehab, Elisa Bertino, and Arif Ghafoor. 2007 · 2007
Earlier work this paper cites.
Fragile database watermarking for malicious tamper detection using support vector regression. In Third International Conference on Intelligent Information Hiding and Multimedia Signal Processing (IIH-MSP 2007) , Vol. 1. IEEE, 493–496
Meng-Hsiun Tsai, Fang-Yu Hsu, Jun-Dong Chang, and Hsien-Chu Wu. 2007 · 2007
Earlier work this paper cites.
A Distortion Free Watermark Framework for Relational Databases.. In ICSOFT (2) . Citeseer, 229–234
Sukriti Bhattacharya, Agostino Cortesi, et al · 2009
Earlier work this paper cites.
Natural language watermarking via morphosyntactic alterations
Hasan Mesut Meral, Bülent Sankur, A Sumru Özsoy, Tunga Güngör, and Emre Sevinç. 2009 · 2009
Earlier work this paper cites.
Ms marco: A human generated machine reading comprehension dataset
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, et al · 2016
Earlier work this paper cites.
A full-text learning to rank dataset for medical information retrieval. In Advances in Information Retrieval: 38th European Conference on IR Research, ECIR 2016, Padua, Italy, March 20–23, 2016. Proceedings 38 . Springer, 716–722
Vera Boteva, Demian Gholipour, Artem Sokolov, and Stefan Riezler. 2016 · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole. 2016 · 2016
Earlier work this paper cites.
Turning your weakness into a strength: Watermarking deep neural networks by backdooring. In 27th USENIX security symposium (USENIX Security 18) . 1615–1631
Yossi Adi, Carsten Baum, Moustapha Cisse, Benny Pinkas, and Joseph Keshet. 2018 · 2018
Earlier work this paper cites.
A comprehensive survey of watermarking relational databases research
Muhammad Kamran and Muddassar Farooq. 2018 · 2018
Earlier work this paper cites.
How much is a triple. In IEEE International Semantic Web Conference
Heiko Paulheim. 2018 · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018 · 2018
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al · 2019
Earlier work this paper cites.
Digital image watermarking techniques: a review
Mahbuba Begum and Mohammad Shorif Uddin. 2020 · 2020
Earlier work this paper cites.
Approximate nearest neighbor negative contrastive learning for dense text retrieval
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul Bennett, Junaid Ahmed, and Arnold Overwijk. 2020 · 2020
Cited alongside, same era.
Unsupervised dense information retrieval with contrastive learning
Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2021 · 2021
Cited alongside, same era.
Towards document-level paraphrase generation with sentence rewriting and reordering
Zhe Lin, Yitao Cai, and Xiaojun Wan. 2021 · 2021
Cited alongside, same era.
TREC-COVID: constructing a pandemic information retrieval test collection. In ACM SIGIR Forum , Vol. 54. ACM New York, NY, USA, 1–12
Ellen Voorhees, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, William R Hersh, Kyle Lo, Kirk Roberts, Ian Soboroff, and Lucy Lu Wang. 2021 · 2021
Cited alongside, same era.
Maya Anderson, Guy Amit, and Abigail Goldsteen. 2024 · 2024
Later among the works it cites.
Undetectable watermarks for language models. In The Thirty Seventh Annual Conference on Learning Theory . PMLR, 1125–1139
Miranda Christ, Sam Gunn, and Or Zamir. 2024 · 2024
Later among the works it cites.
DBpedia
DBpedia Community. 2024 · 2024
Later among the works it cites.
The power of noise: Redefining retrieval for rag systems. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval . 719–729
Florin Cuconasu, Giovanni Trappolini, Federico Siciliano, Simone Filice, Cesare Campagnano, Yoelle Maarek, Nicola Tonellotto, and Fabrizio Silvestri. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022 · 2022
Cited alongside, same era.
Detecting language model attacks with perplexity
Gabriel Alon and Michael Kamfonas. 2023 · 2023
Cited alongside, same era.
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al · 2023
Cited alongside, same era.
Self-rag: Learning to retrieve, generate, and critique through self-reflection
Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2023 · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al · 2023
Cited alongside, same era.
Unbiased watermark for large language models
Zhengmian Hu, Lichang Chen, Xidong Wu, Yihan Wu, Hongyang Zhang, and Heng Huang. 2023 · 2023
Cited alongside, same era.
Baseline defenses for adversarial attacks against aligned language models
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein. 2023 · 2023
Cited alongside, same era.
Active retrieval augmented generation
Zhengbao Jiang, Frank F Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023 · 2023
Cited alongside, same era.
Nikola Jovanović, Robin Staab, Maximilian Baader, and Martin Vechev. 2024 · 2024
Later among the works it cites.
Paraphrasing evades detectors of ai-generated text, but retrieval is an effective defense
Kalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting, and Mohit Iyyer. 2024 · 2024
Later among the works it cites.
Biomedrag: A retrieval augmented large language model for biomedicine
Mingchen Li, Halil Kilicoglu, Hua Xu, and Rui Zhang. 2024a · 2024
Later among the works it cites.
Generating Is Believing: Membership Inference Attacks against Retrieval-Augmented Generation
Yuying Li, Gaoyang Liu, Chen Wang, and Yang Yang. 2024b · 2024
Later among the works it cites.
Ssl-wm: A black-box watermarking approach for encoders pre-trained by self-supervised learning. In Proceedings of the 2024 Annual Network and Distributed System Security Symposium, NDSS’24
Peizhuo Lv, Pan Li, Shenchen Zhu, Shengzhi Zhang, Kai Chen, Ruigang Liang, Chang Yue, Fan Xiang, Yuling Cai, Hualong Ma, et al · 2024
Later among the works it cites.
Revisiting the Robustness of Watermarking to Paraphrasing Attacks
Saksham Rastogi and Danish Pruthi. 2024 · 2024
Later among the works it cites.
Potential for GPT technology to optimize future clinical decision-making using retrieval-augmented generation
Calvin Wang, Joshua Ong, Chara Wang, Hannah Ong, Rebekah Cheng, and Dennis Ong. 2024b · 2024
Later among the works it cites.
Unims-rag: A unified multi-source retrieval-augmented generation for personalized dialogue systems
Hongru Wang, Wenyu Huang, Yang Deng, Rui Wang, Zezhong Wang, Yufei Wang, Fei Mi, Jeff Z Pan, and Kam-Fai Wong. 2024a · 2024
Later among the works it cites.
Statistical Hypothesis Test
Wikipedia. 2024 · 2024
Later among the works it cites.
YAGO Knowledge
YAGO. 2024 · 2024
Later among the works it cites.
Corrective retrieval augmented generation
Shi-Qi Yan, Jia-Chen Gu, Yun Zhu, and Zhen-Hua Ling. 2024 · 2024
Later among the works it cites.
{ \{ REMARK-LLM } \} : A robust and efficient watermarking framework for generative large language models. In 33rd USENIX Security Symposium (USENIX Security 24) . 1813–1830
Ruisi Zhang, Shehzeen Samarah Hussain, Paarth Neekhara, and Farinaz Koushanfar. 2024 · 2024
Later among the works it cites.
Revolutionizing finance with llms: An overview of applications and insights
Huaqin Zhao, Zhengliang Liu, Zihao Wu, Yiwei Li, Tianze Yang, Peng Shu, Shaochen Xu, Haixing Dai, Lin Zhao, Gengchen Mai, et al · 2024
Later among the works it cites.
A survey on generative ai and llm for video generation, understanding, and streaming
Pengyuan Zhou, Lin Wang, Zhi Liu, Yanbin Hao, Pan Hui, Sasu Tarkoma, and Jussi Kangasharju. 2024 · 2024
Later among the works it cites.
Poisonedrag: Knowledge poisoning attacks to retrieval-augmented generation of large language models
Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia. 2024 · 2024
Later among the works it cites.
Code of RAG-WM
2025 · 2025
Closest in time.
Anything LLM AI
Mintplex Labs Inc. 2025 · 2025
Closest in time.
John Snow Labs
John Snow Labs. 2025 · 2025
Closest in time.