Fetching the paper…
Reading the bibliography…
High-quality training data has proven crucial for developing performant large language models (LLMs).
A mathematical theory of communication
Claude Elwood Shannon. 1948 · 1948
Earlier work this paper cites.
Automatic generation of model and data cards: A step towards responsible AI
Jiarui Liu, Wenkai Li, Zhijing Jin, and Mona Diab. 2024b · 1997
Earlier work this paper cites.
A probabilistic Earley parser as a psycholinguistic model
John Hale. 2001 · 2001
Earlier work this paper cites.
Low-level predictive inference in reading: The influence of transitional probabilities on eye movements
Scott A McDonald and Richard C Shillcock. 2003 · 2003
Earlier work this paper cites.
Expectation-based syntactic comprehension
Roger Levy. 2008 · 2008
Earlier work this paper cites.
Membership inference attacks against machine learning models
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. 2017 · 2017
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M. Bender and Batya Friedman. 2018 · 2018
Earlier work this paper cites.
Think you have solved question answering? try ARC, the AI2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018 · 2018
Earlier work this paper cites.
Predictive power of word surprisal for reading times is a linear function of language model quality
Adam Goodkind and Klinton Bicknell. 2018 · 2018
Earlier work this paper cites.
Privacy risk in machine learning: Analyzing the connection to overfitting
Samuel Yeom, Irene Giacomelli, Matt Fredrikson, and Somesh Jha. 2018 · 2018
Earlier work this paper cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019 · 2019
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Úlfar Erlingsson, Alina Oprea, and Colin Raffel. 2021 · 2021
Earlier work this paper cites.
Datasheets for datasets
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2021 · 2021
Earlier work this paper cites.
Quantifying memorization across neural language models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2022 · 2022
Earlier work this paper cites.
Quantifying Privacy Risks of Masked Language Models Using Membership Inference Attacks
Fatemehsadat Mireshghallah, Kartik Goyal, Archit Uniyal, Taylor Berg-Kirkpatrick, and Reza Shokri. 2022 · 2022
Earlier work this paper cites.
Can we trust the evaluation on chatgpt?
Rachith Aiyappa, Jisun An, Haewoon Kwak, and Yong-yeol Ahn. 2023 · 2023
Earlier work this paper cites.
Claude 2
Anthropic. 2023 · 2023
Earlier work this paper cites.
The foundation model transparency index
Rishi Bommasani, Kevin Klyman, Shayne Longpre, Sayash Kapoor, Nestor Maslej, Betty Xiong, Daniel Zhang, and Percy Liang. 2023 · 2023
Earlier work this paper cites.
Speak, memory: An archaeology of books known to ChatGPT/GPT-4
Kent Chang, Mackenzie Cramer, Sandeep Soni, and David Bamman. 2023 · 2023
Earlier work this paper cites.
The chatbot and the canon: Poetry memorization in LLMs
Lyra D’Souza and David Mimno. 2023 · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Team Gemini, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. 2023 · 2023
Cited alongside, same era.
Information value: Measuring utterance predictability as distance from plausible alternatives
Mario Giulianelli, Sarenne Wallbridge, and Raquel Fernández. 2023 · 2023
Cited alongside, same era.
Time travel in LLMs: Tracing data contamination in large language models
Shahriar Golchin and Mihai Surdeanu. 2023 · 2023
Cited alongside, same era.
The Times sues OpenAI and Microsoft over A.I. use of copyrighted work
Michael M Grynbaum and Ryan Mac. 2023 · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Later among the works it cites.
On provable copyright protection for generative models
Nikhil Vyas, Sham M. Kakade, and Boaz Barak. 2023 · 2023
Later among the works it cites.
Don’t make your LLM an evaluation benchmark cheater
Kun Zhou, Yutao Zhu, Zhipeng Chen, Wentong Chen, Wayne Xin Zhao, Xu Chen, Yankai Lin, Ji-Rong Wen, and Jiawei Han. 2023 · 2023
Later among the works it cites.
Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source LLMs
Simone Balloccu, Patrícia Schmidtová, Mateusz Lango, and Ondrej Dusek. 2024 · 2024
Later among the works it cites.
Copyright and ai training data—transparency to the rescue?
Adam Buick. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Preventing generation of verbatim memorization in language models gives a false sense of privacy
Daphne Ippolito, Florian Tramer, Milad Nasr, Chiyuan Zhang, Matthew Jagielski, Katherine Lee, Christopher Choquette Choo, and Nicholas Carlini. 2023 · 2023
Cited alongside, same era.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023 · 2023
Cited alongside, same era.
Copyright violations and large language models
Antonia Karamolegkou, Jiaang Li, Li Zhou, and Anders Søgaard. 2023a · 2023
Cited alongside, same era.
Copyright violations and large language models
Antonia Karamolegkou, Jiaang Li, Li Zhou, and Anders Søgaard. 2023b · 2023
Cited alongside, same era.
LLM360: Towards fully transparent open-source LLMs
Zhengzhong Liu, Aurick Qiao, Willie Neiswanger, Hongyi Wang, Bowen Tan, Tianhua Tao, Junbo Li, Yuqi Wang, Suqi Sun, Omkar Pangarkar, et al. 2023 · 2023
Cited alongside, same era.
Membership inference attacks against language models via neighbourhood comparison
Justus Mattern, Fatemehsadat Mireshghallah, Zhijing Jin, Bernhard Schoelkopf, Mrinmaya Sachan, and Taylor Berg-Kirkpatrick. 2023 · 2023
Cited alongside, same era.
Silo language models: Isolating legal risk in a nonparametric datastore
Sewon Min, Suchin Gururangan, Eric Wallace, Hannaneh Hajishirzi, Noah A Smith, and Luke Zettlemoyer. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
Investigating data contamination in modern benchmarks for large language models
Chunyuan Deng, Yilun Zhao, Xiangru Tang, Mark Gerstein, and Arman Cohan. 2024 · 2024
Later among the works it cites.
Do membership inference attacks work on large language models?
Michael Duan, Anshuman Suri, Niloofar Mireshghallah, Sewon Min, Weijia Shi, Luke Zettlemoyer, Yulia Tsvetkov, Yejin Choi, David Evans, and Hannaneh Hajishirzi. 2024 · 2024
Later among the works it cites.
De-cop: Detecting copyrighted content in language models training data
André V Duarte, Xuandong Zhao, Arlindo L Oliveira, and Lei Li. 2024 · 2024
Later among the works it cites.
What’s in my big data?
Yanai Elazar, Akshita Bhagia, Ian Helgi Magnusson, Abhilasha Ravichander, Dustin Schwenk, Alane Suhr, Evan Pete Walsh, Dirk Groeneveld, Luca Soldaini, Sameer Singh, Hanna Hajishirzi, Noah A. Smith, and Jesse Dodge. 2024 · 2024
Later among the works it cites.
Investigating data contamination for pre-training language models
Minhao Jiang, Ken Ziyu Liu, Ming Zhong, Rylan Schaeffer, Siru Ouyang, Jiawei Han, and Sanmi Koyejo. 2024 · 2024
Later among the works it cites.
Data authenticity, consent, and provenance for ai are all broken: What will it take to fix them?
Shayne Longpre, Robert Mahari, Naana Obeng-Marnu, William Brannon, Tobin South, Jad Kabbara, and Sandy Pentland. 2024 · 2024
Later among the works it cites.
LLM dataset inference: Did you train on my dataset?
Pratyush Maini, Hengrui Jia, Nicolas Papernot, and Adam Dziedzic. 2024 · 2024
Later among the works it cites.
Copyright traps for large language models
Matthieu Meeus, Igor Shilov, Manuel Faysse, and Yves-Alexandre De Montjoye. 2024 · 2024
Later among the works it cites.
Data contamination report from the 2024 CONDA shared task
Oscar Sainz, Iker García-Ferrero, Alon Jacovi, Jon Ander Campos, Yanai Elazar, Eneko Agirre, Yoav Goldberg, Wei-Lin Chen, Jenny Chim, Leshem Choshen, Luca D’Amico-Wong, Melissa Dell, Run-Ze Fan, Shahriar Golchin, Yucheng Li, Pengfei Liu, Bhavish Pahwa, Ameya Prabhu, Suryansh Sharma, Emily Silcock, Kateryna Solonko, David Stap, Mihai Surdeanu, Yu-Min Tseng, Vishaal Udandarao, Zengzhi Wang, Ruijie Xu, and Jinglin Yang. 2024 · 2024
Later among the works it cites.
Artificial intelligence and privacy
Daniel J Solove. 2024 · 2024
Later among the works it cites.
Sonnet or not, bot? poetry evaluation for large models and datasets
Melanie Walsh, Maria Antoniak, and Anna Preus. 2024 · 2024
Later among the works it cites.
“According to . . . ”: Prompting language models improves quoting from pre-training data
Orion Weller, Marc Marone, Nathaniel Weir, Dawn Lawrie, Daniel Khashabi, and Benjamin Van Durme. 2024 · 2024
Later among the works it cites.
Benchmark data contamination of large language models: A survey
Cheng Xu, Shuhao Guan, Derek Greene, M Kechadi, et al. 2024 · 2024
Later among the works it cites.