Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have transformed machine learning but raised significant legal concerns due to their potential to produce text that infringes on copyrights, resulting in several high-profile lawsuits.
Berne Convention for the Protection of Literary and Artistic Works
World Intellectual Property Organization (WIPO). 1971 · 1971
Earlier work this paper cites.
Internet Archive: Digital Library
Internet Archive. 1996 · 1996
Earlier work this paper cites.
Google Books: Search and Preview Books
Google Books. 2004 · 2004
Earlier work this paper cites.
ManyBooks: Free eBooks
ManyBooks. 2004 · 2004
Earlier work this paper cites.
LibriVox: Free Public Domain Audiobooks
LibriVox. 2005 · 2005
Earlier work this paper cites.
Open Library: An Open, Editable Library Catalog
Open Library. 2006 · 2006
Earlier work this paper cites.
HathiTrust Digital Library
HathiTrust. 2008 · 2008
Earlier work this paper cites.
Understanding Copyright and Related Rights
World Intellectual Property Organization. 2016 · 2016
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al. 2021 · 2021
Earlier work this paper cites.
Quantifying memorization across neural language models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2022 · 2022
Earlier work this paper cites.
Speak, memory: An archaeology of books known to chatgpt/gpt-4
Kent Chang, Mackenzie Cramer, Sandeep Soni, and David Bamman. 2023 · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong. 2023 · 2023
Earlier work this paper cites.
Unlearn what you want to forget: Efficient unlearning for llms
Jiaao Chen and Diyi Yang. 2023 · 2023
Earlier work this paper cites.
The chatbot and the canon: Poetry memorization in llms
Lyra D’Souza and David Mimno. 2023 · 2023
Earlier work this paper cites.
Who’s harry potter? approximate unlearning in llms
Ronen Eldan and Mark Russinovich. 2023 · 2023
Earlier work this paper cites.
Foundation models and fair use
Peter Henderson, Xuechen Li, Dan Jurafsky, Tatsunori Hashimoto, Mark A Lemley, and Percy Liang. 2023 · 2023
Earlier work this paper cites.
Preventing generation of verbatim memorization in language models gives a false sense of privacy
Daphne Ippolito, Florian Tramer, Milad Nasr, Chiyuan Zhang, Matthew Jagielski, Katherine Lee, Christopher Choquette Choo, and Nicholas Carlini. 2023 · 2023
Earlier work this paper cites.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023 · 2023
Earlier work this paper cites.
Copyright violations and large language models
Antonia Karamolegkou, Jiaang Li, Li Zhou, and Anders Søgaard. 2023 · 2023
Earlier work this paper cites.
Deepinception: Hypnotize large language model to be jailbreaker
Xuan Li, Zhanke Zhou, Jianing Zhu, Jiangchao Yao, Tongliang Liu, and Bo Han. 2023 · 2023
Earlier work this paper cites.
Unsupervised entity alignment for temporal knowledge graphs
Xiaoze Liu, Junyang Wu, Tianyi Li, Lu Chen, and Yunjun Gao. 2023b · 2023
Cited alongside, same era.
Silo language models: Isolating legal risk in a nonparametric datastore
Sewon Min, Suchin Gururangan, Eric Wallace, Hannaneh Hajishirzi, Noah A Smith, and Luke Zettlemoyer. 2023 · 2023
Cited alongside, same era.
Scalable extraction of training data from (production) language models
Milad Nasr, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A. Feder Cooper, Daphne Ippolito, Christopher A. Choquette-Choo, Eric Wallace, Florian Tramèr, and Katherine Lee. 2023 · 2023
Cited alongside, same era.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. 2023 · 2023
Cited alongside, same era.
Be like a goldfish, don’t memorize! mitigating memorization in generative llms
Abhimanyu Hans, Yuxin Wen, Neel Jain, John Kirchenbauer, Hamid Kazemi, Prajwal Singhania, Siddharth Singh, Gowthami Somepalli, Jonas Geiping, Abhinav Bhatele, et al. 2024 · 2024
Closest in time.
Digger: Detecting copyright content mis-usage in large language model training
Haodong Li, Gelei Deng, Yi Liu, Kailong Wang, Yuekang Li, Tianwei Zhang, Yang Liu, Guoai Xu, Guosheng Xu, and Haoyu Wang. 2024 · 2024
Closest in time.
Prominent authors sue openai over chatbot technology
Sapna Maheshwari and Marc Tracy. 2023 · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date
Meta. 2024 · 2024
Closest in time.
Llms and memorization: On quality and specificity of copyright compliance
Felix B Mueller, Rebekka Görge, Anna K Bernzen, Janna C Pirk, and Maximilian Poretschkin. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tianhao Shen, Renren Jin, Yufei Huang, Chuang Liu, Weilong Dong, Zishan Guo, Xinwei Wu, Yan Liu, and Deyi Xiong. 2023 · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. 2023 · 2023
Cited alongside, same era.
Survey on factuality in large language models: Knowledge, retrieval and domain-specificity
Cunxiang Wang, Xiaoze Liu, Yuanhao Yue, Xiangru Tang, Tianhang Zhang, Cheng Jiayang, Yunzhi Yao, Wenyang Gao, Xuming Hu, Zehan Qi, et al. 2023 · 2023
Cited alongside, same era.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023 · 2023
Cited alongside, same era.
Large language model unlearning
Yuanshun Yao, Xiaojun Xu, and Yang Liu. 2023 · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. 2023 · 2023
Cited alongside, same era.
Strong copyright protection for language models via adaptive model fusion
Javier Abad, Konstantin Donhauser, Francesco Pinto, and Fanny Yang. 2024 · 2024
Cited alongside, same era.
Sarah silverman sues meta and openai
Abigail Adams. 2023 · 2024
Cited alongside, same era.
Closest in time.
Meet dan: The jailbreak version of chatgpt and how to use it - ai unchained and unfiltered
Neonforge. 2023 · 2024
Closest in time.
How long does copyright protection last?
U.S. Copyright Office. 2023 · 2024
Closest in time.
Hello gpt-4o
OpenAI. 2024a · 2024
Closest in time.
Introducing chatgpt and whisper apis
OpenAI. 2024b · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al. 2024 · 2024
Closest in time.
Rethinking llm memorization through the lens of adversarial compression
Avi Schwarzschild, Zhili Feng, Pratyush Maini, Zachary C. Lipton, and J. Zico Kolter. 2024 · 2024
Closest in time.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. 2024 · 2024
Closest in time.
Welcome to the public domain
Rich Stim. 2013 · 2024
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, et al. 2023 · 2024
Closest in time.
The new york times sues openai and microsoft over copyright infringement
Marc Tracy and Sapna Maheshwari. 2023 · 2024
Closest in time.
Copyright renewals database
Stanford University. 2023 · 2024
Closest in time.
Evaluating copyright takedown methods for language models
Boyi Wei, Weijia Shi, Yangsibo Huang, Noah A Smith, Chiyuan Zhang, Luke Zettlemoyer, Kai Li, and Peter Henderson. 2024 · 2024
Closest in time.
List of most-streamed songs on spotify — wikipedia, the free encyclopedia
Wikipedia. 2024 · 2024
Closest in time.
Safedecoding: Defending against jailbreak attacks via safety-aware decoding
Zhangchen Xu, Fengqing Jiang, Luyao Niu, Jinyuan Jia, Bill Yuchen Lin, and Radha Poovendran. 2024 · 2024
Closest in time.