Fetching the paper…
Reading the bibliography…
Frontier AI systems are making transformative impacts across society, but such benefits are not without costs: models trained on web-scale datasets containing personal and private data raise profound concerns about data privacy and security.
Sorami Hisamoto, Matt Post and Kevin Duh · 1904
Earlier work this paper cites.
“A continual learning survey: Defying forgetting in classification tasks”, 2019
Matthias De et al · 1909
Earlier work this paper cites.
“Analyzing Information Leakage of Updates to Natural Language Models”, 2019
Santiago Zanella-Béguelin et al · 1912
Earlier work this paper cites.
“On tables of random numbers”
Andrei Kolmogorov · 1963
Earlier work this paper cites.
“Binary codes capable of correcting deletions, insertions, and reversals”
Vladimir Levenshtein · 1965
Earlier work this paper cites.
“On tables of random numbers”
Andrei Kolmogorov · 1998
Earlier work this paper cites.
“REALM: Retrieval-Augmented Language Model Pre-Training”, 2020
Kelvin Guu et al · 2002
Earlier work this paper cites.
“Recall and Learn: Fine-tuning Deep Pretrained Language Models with Less Forgetting”, 2020
Sanyuan Chen et al · 2004
Earlier work this paper cites.
“Language Models are Few-Shot Learners”, 2020
Tom Brown et al · 2005
Earlier work this paper cites.
“Understanding Unintended Memorization in Federated Learning”, 2020
Om Thakkar, Swaroop Ramaswamy, Rajiv Mathews and Françoise Beaufays · 2006
Earlier work this paper cites.
“What Neural Networks Memorize and Why: Discovering the Long Tail via Influence Estimation”, 2020
Vitaly Feldman and Chiyuan Zhang · 2008
Earlier work this paper cites.
“Extracting Training Data from Large Language Models”, 2020
Nicholas Carlini et al · 2012
Earlier work this paper cites.
“Membership Inference Attacks against Machine Learning Models”, 2016
Reza Shokri, Marco Stronati, Congzheng Song and Vitaly Shmatikov · 2016
Earlier work this paper cites.
“Ethical Challenges in Data-Driven Dialogue Systems”, 2017
Peter Henderson et al · 2017
Earlier work this paper cites.
“Overcoming catastrophic forgetting in neural networks”
James Kirkpatrick et al · 2017
Earlier work this paper cites.
“Continual Learning Through Synaptic Intelligence”
Friedemann Zenke, Ben Poole and Surya Ganguli · 2017
Earlier work this paper cites.
“Online Structured Laplace Approximations For Overcoming Catastrophic Forgetting”, 2018
Hippolyt Ritter, Aleksandar Botev and David Barber · 2018
Cited alongside, same era.
“Progress & Compress: A scalable framework for continual learning”, 2018
Jonathan Schwarz et al · 2018
Cited alongside, same era.
“Overcoming catastrophic forgetting with hard attention to the task”, 2018
Joan Serrà, Dídac Surís, Marius Miron and Alexandros Karatzoglou · 2018
Cited alongside, same era.
“An Empirical Study of Example Forgetting during Deep Neural Network Learning”, 2018
Mariya Toneva et al · 2018
Cited alongside, same era.
“The Pile: An 800GB dataset of diverse text for language modeling”
“Gemini: A Family of Highly Capable Multimodal Models”, 2023
Gemini Team et al · 2023
Later among the works it cites.
“Copyright Violations and Large Language Models”, 2023
Antonia Karamolegkou, Jiaang Li, Li Zhou and Anders Sogaard · 2023
Later among the works it cites.
“LLM360: Towards Fully Transparent Open-Source LLMs”, 2023
Zhengzhong Liu et al · 2023
Later among the works it cites.
“GPT-4 Technical Report”, 2023
OpenAI et al · 2023
Later among the works it cites.
“Are Emergent Abilities of Large Language Models a Mirage?”
Rylan Schaeffer, Brando Miranda and Sanmi Koyejo · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Leo Gao et al · 2020
Cited alongside, same era.
“Deduplicating Training Data Makes Language Models Better”, 2021
Katherine Lee et al · 2021
Cited alongside, same era.
“Wide Neural Networks Forget Less Catastrophically”, 2021
Seyed Mirzadeh et al · 2021
Cited alongside, same era.
“Counterfactual Memorization in Neural Language Models”, 2021
Chiyuan Zhang et al · 2021
Cited alongside, same era.
“A Review on Language Models as Knowledge Bases”, 2022
Badr AlKhamissi et al · 2022
Cited alongside, same era.
“Quantifying Memorization Across Neural Language Models”, 2022
Nicholas Carlini et al · 2022
Cited alongside, same era.
“Deduplicating Training Data Mitigates Privacy Risks in Language Models”, 2022
Nikhil Kandpal, Eric Wallace and Colin Raffel · 2022
Cited alongside, same era.
“Quantifying Privacy Risks of Masked Language Models Using Membership Inference Attacks”, 2022
Fatemehsadat Mireshghallah et al · 2022
Cited alongside, same era.
“Detecting Pretraining Data from Large Language Models”, 2023
Weijia Shi et al · 2023
Later among the works it cites.
“Identifying and Mitigating Privacy Risks Stemming from Language Models: A Survey”, 2023
Victoria Smith, Ali Shamsabadi, Carolyn Ashurst and Adrian Weller · 2023
Later among the works it cites.
“Beyond Memorization: Violating Privacy Via Inference with Large Language Models”, 2023
Robin Staab, Mark Vero, Mislav Balunović and Martin Vechev · 2023
Later among the works it cites.
“Assessing Privacy Risks in Language Models: A Case Study on Summarization Tasks”, 2023
Ruixiang Tang et al · 2023
Later among the works it cites.
“LLaMA: Open and efficient foundation language models”
Hugo Touvron et al · 2023
Later among the works it cites.
“A Comprehensive Survey of Continual Learning: Theory, Method and Application”, 2023
Liyuan Wang, Xingxing Zhang, Hang Su and Jun Zhu · 2023
Later among the works it cites.
“Elephants Never Forget: Memorization and Learning of Tabular Data in Large Language Models”, 2024
Sebastian Bordt et al · 2024
Closest in time.
“Do Membership Inference Attacks Work on Large Language Models?”, 2024
Michael Duan et al · 2024
Closest in time.
“Alpaca against Vicuna: Using LLMs to Uncover Memorization of LLMs”, 2024
Aly Kassem et al · 2024
Closest in time.
“gzip Predicts Data-dependent Scaling Laws”, 2024
Rohan Pandey · 2024
Closest in time.
Rylan Schaeffer et al · 2024
Closest in time.