Fetching the paper…
Reading the bibliography…
RAG enables LLMs to easily incorporate external data, raising concerns for data owners regarding unauthorized usage of their content.
The enron corpus: A new dataset for email classification research
Bryan Klimt and Yiming Yang · 2004
Earlier work this paper cites.
Membership inference attacks against machine learning models
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov · 2017
Earlier work this paper cites.
Turning your weakness into a strength: Watermarking deep neural networks by backdooring
Yossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas, and Joseph Keshet · 2018
Earlier work this paper cites.
Directive (eu) 2019/790 of the european parliament and of the council on copyright and related rights in the digital single market and amending directives 96/9/ec and 2001/29/ec
Council of the European Union · 2019
Earlier work this paper cites.
Proposal for a regulation of the european parliament and of the council laying down harmonised rules on artificial intelligence (artificial intelligence act) and amending certain union legislative acts - analysis of the final compromise text with a view to agreement
European Parliament and the Council of the European Union · 2019
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive NLP tasks
Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela · 2020
Earlier work this paper cites.
MedDialog: Large-scale medical dialogue datasets
Guangtao Zeng, Wenmian Yang, Zeqian Ju, Yue Yang, Sicheng Wang, Ruisi Zhang, Meng Zhou, Jiaqi Zeng, Xiangyu Dong, Ruoyu Zhang, Hongchao Fang, Penghui Zhu, Shu Chen, and Pengtao Xie · 2020
Earlier work this paper cites.
Bertscore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi · 2020
Earlier work this paper cites.
Label-only membership inference attacks
Christopher A. Choquette-Choo, Florian Tramèr, Nicholas Carlini, and Nicolas Papernot · 2021
Earlier work this paper cites.
Technical specification v1.3. technical report. c2pa., 2021
Coalition for Content Provenance and Authenticity · 2021
Earlier work this paper cites.
Dataset inference: Ownership resolution in machine learning
Pratyush Maini, Mohammad Yaghini, and Nicolas Papernot · 2021
Earlier work this paper cites.
Joint summarization-entailment optimization for consumer health question understanding
Khalil Mrini, Franck Dernoncourt, Walter Chang, Emilia Farcas, and Ndapandula Nakashole · 2021
Earlier work this paper cites.
Membership inference attacks from first principles
Nicholas Carlini, Steve Chien, Milad Nasr, Shuang Song, Andreas Terzis, and Florian Tramèr · 2022
Earlier work this paper cites.
Dataset inference for self-supervised models
Adam Dziedzic, Haonan Duan, Muhammad Ahmad Kaleem, Nikita Dhawan, Jonas Guan, Yannis Cattan, Franziska Boenisch, and Nicolas Papernot · 2022
Earlier work this paper cites.
Paraphrastic representations at scale
John Wieting, Kevin Gimpel, Graham Neubig, and Taylor Berg-Kirkpatrick · 2022
Earlier work this paper cites.
Decorait – decentralized opt-in/out registry for ai training, 2023
Kar Balan, Alex Black, Simon Jenni, Andrew Gilbert, Andy Parsons, and John Collomosse · 2023
Earlier work this paper cites.
Flocks of stochastic parrots: Differentially private prompt learning for large language models
Haonan Duan, Adam Dziedzic, Nicolas Papernot, and Franziska Boenisch · 2023
Earlier work this paper cites.
Retrieval-augmented generation for large language models: A survey
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Qianyu Guo, Meng Wang, and Haofen Wang · 2023
Earlier work this paper cites.
The times sues openai and microsoft over a.i. use of copyrighted work, 2023
Michael M. Grynbaum and Ryan Mac · 2023
Earlier work this paper cites.
Domain watermark: Effective and harmless dataset copyright protection is closed at hand
Junfeng Guo, Yiming Li, Lixu Wang, Shu-Tao Xia, Heng Huang, Cong Liu, and Bo Li · 2023
Earlier work this paper cites.
Preventing generation of verbatim memorization in language models gives a false sense of privacy
Daphne Ippolito, Florian Tramèr, Milad Nasr, Chiyuan Zhang, Matthew Jagielski, Katherine Lee, Christopher A. Choquette-Choo, and Nicholas Carlini · 2023
Cited alongside, same era.
Defining best practices for opting out of ml training
Paul Keller and Zuzanna Warso · 2023
Cited alongside, same era.
A watermark for large language models
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein · 2023
Cited alongside, same era.
The data provenance initiative: A large scale audit of dataset licensing & attribution in ai
Shayne Longpre, Robert Mahari, Anthony Chen, Naana Obeng-Marnu, Damien Sileo, William Brannon, Niklas Muennighoff, Nathan Khazam, Jad Kabbara, Kartik Perisetla, Xinyi Wu, Enrico Shippole, Kurt Bollacker, Tongshuang Wu, Luis Villa, Sandy Pentland, and Sara Hooker · 2023
Cited alongside, same era.
Mark my words: Analyzing and evaluating language model watermarks
Julien Piet, Chawin Sitawarin, Vivian Fang, Norman Mu, and David A. Wagner · 2023
Cited alongside, same era.
Robust distortion-free watermarks for language models
Rohith Kuditipudi, John Thickstun, Tatsunori Hashimoto, and Percy Liang · 2024
Closest in time.
Seeing is believing: Black-box membership inference attacks against retrieval augmented generation
Yuying Li, Gaoyang Liu, Yang Yang, and Chen Wang · 2024
Closest in time.
A private watermark for large language models
Aiwei Liu, Leyi Pan, Xuming Hu, Shuang Li, Lijie Wen, Irwin King, and Philip S. Yu · 2024
Closest in time.
LLM dataset inference: Did you train on my dataset?
Pratyush Maini, Hengrui Jia, Nicolas Papernot, and Adam Dziedzic · 2024
Closest in time.
Inherent challenges of post-hoc membership inference for large language models
Matthieu Meeus, Shubham Jain, Marek Rei, and Yves-Alexandre de Montjoye · 2024
Closest in time.
Repliqa: A question-answering dataset for benchmarking llms on unseen reference content
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ai.txt, 2023
SpawningAI · 2023
Cited alongside, same era.
Protecting language generation models via invisible watermarking
Xuandong Zhao, Yu-Xiang Wang, and Lei Li · 2023
Cited alongside, same era.
Is my data in your retrieval database? membership inference attacks against retrieval augmented generation
Maya Anderson, Guy Amit, and Abigail Goldsteen · 2024
Cited alongside, same era.
Stability ai, midjourney should face artists’ copyright case, judge says, 2024
Blake Brittain · 2024
Cited alongside, same era.
Phantom: General trigger attacks on retrieval augmented language generation
Harsh Chaudhari, Giorgio Severi, John Abascal, Matthew Jagielski, Christopher A. Choquette-Choo, Milad Nasr, Cristina Nita-Rotaru, and Alina Oprea · 2024
Cited alongside, same era.
Undetectable watermarks for language models
Miranda Christ, Sam Gunn, and Or Zamir · 2024
Cited alongside, same era.
Blind baselines beat membership inference attacks for foundation models
Debeshee Das, Jie Zhang, and Florian Tramèr · 2024
Cited alongside, same era.
Joao Monteiro, Pierre-Andre Noel, Etienne Marcotte, Sai Rajeswar, Valentina Zantedeschi, David Vazquez, Nicolas Chapados, Christopher Pal, and Perouz Taslakian · 2024
Closest in time.
Proving test set contamination in black-box language models
Yonatan Oren, Nicole Meister, Niladri S. Chatterji, Faisal Ladhak, and Tatsunori Hashimoto · 2024
Closest in time.
Entruth: Enhancing the traceability of unauthorized dataset usage in text-to-image diffusion models with minimal and robust alterations
Jie Ren, Yingqian Cui, Chen Chen, Vikash Sehwag, Yue Xing, Jiliang Tang, and Lingjuan Lyu · 2024
Closest in time.
Watermarking makes language models radioactive
Tom Sander, Pierre Fernandez, Alain Durmus, Matthijs Douze, and Teddy Furon · 2024
Closest in time.
Spawningai, 2024
SpawningAI · 2024
Closest in time.
Tdm reservation protocol (tdmrep) - final community group report, 2024
TDM Community Group · 2024
Closest in time.
Badrag: Identifying vulnerabilities in retrieval augmented generation of large language models
Jiaqi Xue, Mengxin Zheng, Yebowen Hu, Fei Liu, Xun Chen, and Qian Lou · 2024
Closest in time.
The good and the bad: Exploring privacy issues in retrieval-augmented generation (rag)
Shenglai Zeng, Jiankun Zhang, Pengfei He, Yue Xing, Yiding Liu, Han Xu, Jie Ren, Shuaiqiang Wang, Dawei Yin, Yi Chang, et al · 2024
Closest in time.
Provable robust watermarking for ai-generated text
Xuandong Zhao, Prabhanjan Ananth, Lei Li, and Yu-Xiang Wang · 2024
Closest in time.
Kneschke vs. laion - landmark ruling on tdm exceptions for ai training data, 2024
Paul Goldstein, Christiane Stuetzle, and Susan Bischoff · 2025
Closest in time.
The devil is in the training data, 2023
A. Feder Cooper Katherine Lee, Daphne Ippolito · 2025
Closest in time.
Generative ai, copyright and the ai act, 2023
Joao Pedro Quintais · 2025
Closest in time.
Have i been trained?, 2022
SpawningAI · 2025
Closest in time.
Poisonedrag: Knowledge poisoning attacks to retrieval-augmented generation of large language models
Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia · 2025
Closest in time.