Fetching the paper…
Reading the bibliography…
Exploring the data sources used to train Large Language Models (LLMs) is a crucial direction in investigating potential copyright infringement by these models.
Automatic evaluation of summaries using n-gram co-occurrence statistics. In Proceedings of the 2003 human language technology conference of the North American chapter of the association for computational linguistics . 150–157
Chin-Yew Lin and Eduard Hovy. 2003 · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries. In Text summarization branches out . 74–81
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics. In Proceedings of the 42nd annual meeting of the association for computational linguistics (ACL-04) . 605–612
Chin-Yew Lin and Franz Josef Och. 2004 · 2004
Earlier work this paper cites.
Copyright Law of the United States and Related Laws Contained in Tıtle 17 of the United States Code
U.S. Copyright Office. 1976a · 2022
Earlier work this paper cites.
Copyright Law of the United States and Related Laws Contained in Tıtle 17 of the United States Code
U.S. Copyright Office. 1976b · 2022
Earlier work this paper cites.
Copyright violations and large language models
Antonia Karamolegkou, Jiaang Li, Li Zhou, and Anders Søgaard. 2023 · 2023
Earlier work this paper cites.
Andersen v. Stability AI Ltd
Ethical Tech Initiative of George Washington Law School. 2023a · 2023
Earlier work this paper cites.
GPT-4 Architecture, Infrastructure, Training Dataset, Costs, Vision, MoE
DYLAN PATEL. 2023 · 2023
Cited alongside, same era.
Beyond fair use: Legal risk evaluation for training LLMs on copyrighted text. In ICML Workshop on Generative AI and Law
Noorjahan Rahman and Eduardo Santacana. 2023 · 2023
Cited alongside, same era.
Detecting pretraining data from large language models
Weijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang, Daogao Liu, Terra Blevins, Danqi Chen, and Luke Zettlemoyer. 2023 · 2023
Cited alongside, same era.
Tong Chen, Akari Asai, Niloofar Mireshghallah, Sewon Min, James Grimmelmann, Yejin Choi, Hannaneh Hajishirzi, Luke Zettlemoyer, and Pang Wei Koh. 2024 · 2024
Cited alongside, same era.
Tremblay v. OpenAI, Inc. (In re OpenAI ChatGPT Litigation)
Ethical Tech Initiative of George Washington Law School. 2023b · 2024
Closest in time.
The role of llms in sustainable smart cities: Applications, challenges, and future directions
Amin Ullah, Guilin Qi, Saddam Hussain, Irfan Ullah, and Zafar Ali. 2024 · 2024
Closest in time.
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al · 2024
Closest in time.
A survey on large language model (LLM) security and privacy: The Good, The Bad, and The Ugly
Yifan Yao, Jinhao Duan, Kaidi Xu, Yuanfang Cai, Zhibo Sun, and Yue Zhang. 2024 · 2024
Closest in time.
Yi: Open foundation models by 01. ai
Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Matthieu Meeus, Igor Shilov, Manuel Faysse, and Yves-Alexandre de Montjoye. 2024b · 2024
Cited alongside, same era.
LLMs and Memorization: On Quality and Specificity of Copyright Compliance
Felix B Mueller, Rebekka Görge, Anna K Bernzen, Janna C Pirk, and Maximilian Poretschkin. 2024 · 2024
Cited alongside, same era.
Did the neurons read your book? document-level membership inference for large language models. In 33rd USENIX Security Symposium (USENIX Security 24) . 2369–2385
Matthieu Meeus, Shubham Jain, Marek Rei, and Yves-Alexandre de Montjoye. 2024a
Cited in the paper.
Closest in time.
Revolutionizing finance with llms: An overview of applications and insights
Huaqin Zhao, Zhengliang Liu, Zihao Wu, Yiwei Li, Tianze Yang, Peng Shu, Shaochen Xu, Haixing Dai, Lin Zhao, Gengchen Mai, et al · 2024
Closest in time.