Fetching the paper…
Reading the bibliography…
Large language models are susceptible to memorizing repeated sequences, posing privacy and copyright concerns.
Improving generalization by controlling label-noise information in neural network weights, 2020
Hrayr Harutyunyan, Kyle Reing, Greg Ver Steeg, and Aram Galstyan · 2002
Earlier work this paper cites.
What neural networks memorize and why: Discovering the long tail via influence estimation, 2020
Vitaly Feldman and Chiyuan Zhang · 2008
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer, 2017
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
An empirical study of example forgetting during deep neural network learning
Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J Gordon · 2018
Earlier work this paper cites.
Satrajit Chatterjee · 2020
Earlier work this paper cites.
Gradient descent with early stopping is provably robust to label noise for overparameterized neural networks
Mingchen Li, Mahdi Soltanolkotabi, and Samet Oymak · 2020
Earlier work this paper cites.
Machine unlearning
Lucas Bourtoule, Varun Chandrasekaran, Christopher A Choquette-Choo, Hengrui Jia, Adelin Travers, Baiwu Zhang, David Lie, and Nicolas Papernot · 2021
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy · 2021
Earlier work this paper cites.
Demix layers: Disentangling domains for modular language modeling, 2021
Suchin Gururangan, Mike Lewis, Ari Holtzman, Noah A. Smith, and Luke Zettlemoyer · 2021
Earlier work this paper cites.
Continual learning and private unlearning, 2022
Bo Liu, Qiang Liu, and Peter Stone · 2022
Earlier work this paper cites.
Characterizing datapoints via second-split forgetting
Pratyush Maini, Saurabh Garg, Zachary Lipton, and J. Zico Kolter · 2022
Cited alongside, same era.
Unrolling sgd: Understanding factors influencing machine unlearning, 2022
Anvith Thudi, Gabriel Deza, Varun Chandrasekaran, and Nicolas Papernot · 2022
Cited alongside, same era.
St-moe: Designing stable and transferable sparse expert models, 2022
Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus · 2022
Cited alongside, same era.
Quantifying memorization across neural language models, 2023
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang · 2023
Cited alongside, same era.
Tinystories: How small can language models be and still speak coherent english?, 2023
George-Octavian Barbulescu and Peter Triantafillou · 2024
Later among the works it cites.
Learnable privacy neurons localization in language models, 2024
Ruizhe Chen, Tianxiang Hu, Yang Feng, and Zuozhu Liu · 2024
Later among the works it cites.
Gradient routing: Masking gradients to localize computation in neural networks, 2024
Alex Cloud, Jacob Goldman-Wetzler, Evžen Wybitul, Joseph Miller, and Alexander Matt Turner · 2024
Later among the works it cites.
Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models, 2024
Damai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y. Wu, Zhenda Xie, Y. K. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, and Wenfeng Liang · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ronen Eldan and Yuanzhi Li · 2023
Cited alongside, same era.
Can neural network memorization be localized?, 2023
Pratyush Maini, Michael C. Mozer, Hanie Sedghi, Zachary C. Lipton, J. Zico Kolter, and Chiyuan Zhang · 2023
Cited alongside, same era.
Fact finding: Attempting to reverse-engineer factual recall on the neuron level, Dec 2023
Neel Nanda, Senthooran Rajamanoharan, Janos Kramar, and Rohin Shah · 2023
Cited alongside, same era.
Scalable extraction of training data from (production) language models, 2023
Milad Nasr, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A. Feder Cooper, Daphne Ippolito, Christopher A. Choquette-Choo, Eric Wallace, Florian Tramèr, and Katherine Lee · 2023
Cited alongside, same era.
Neurips 2023 - machine unlearning
Eleni Triantafillou, Fabian Pedregosa, Jamie Hayes, Peter Kairouz, Isabelle Guyon, Meghdad Kurmanji, Gintare Karolina Dziugaite, Peter Triantafillou, Kairan Zhao, Lisheng Sun Hosoya, Julio C. S. Jacques Junior, Vincent Dumoulin, Ioannis Mitliagkas, Sergio Escalera, Jun Wan, Sohier Dane, Maggie Demkin, and Walter Reade · 2023
Cited alongside, same era.
On the spectral bias of two-layer linear networks
Aditya Vardhan Varre, Maria-Luiza Vladarean, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2023
Cited alongside, same era.
How do large language models acquire factual knowledge during pretraining?, 2024a
Hoyeon Chang, Jinho Park, Seonghyeon Ye, Sohee Yang, Youngkyung Seo, Du-Seong Chang, and Minjoon Seo
Cited in the paper.
Do localization methods actually localize memorized data in llms? a tale of two benchmarks, 2024b
Ting-Yun Chang, Jesse Thomason, and Robin Jia
Cited in the paper.
Zhiqiang Shen, Tianhua Tao, Liqun Ma, Willie Neiswanger, Zhengzhong Liu, Hongyi Wang, Bowen Tan, Joel Hestness, Natalia Vassilieva, Daria Soboleva, and Eric Xing · 2024
Later among the works it cites.
Localizing paragraph memorization in language models, 2024
Niklas Stoehr, Mitchell Gordon, Chiyuan Zhang, and Owen Lewis · 2024
Later among the works it cites.
Negative preference optimization: From catastrophic collapse to effective unlearning, 2024
Ruiqi Zhang, Licong Lin, Yu Bai, and Song Mei · 2024
Later among the works it cites.
huggingface/smollm
Loubna Ben Allal, elie, Andrés Marafioti, Anton Lozhkov, Quentin Lhoest, Merve Noyan, Miquel Farré, Gabriel Martín Blázquez, Nouamane Tazi, yousan, Alvaro Bartolome, Emmanuel Ferdman, Jafar Isbarov, vb, and Lionel Cheng · 2025
Closest in time.
Monet: Mixture of monosemantic experts for transformers, 2025
Jungwoo Park, Young Jin Ahn, Kee-Eung Kim, and Jaewoo Kang · 2025
Closest in time.
Polysemanticity and capacity in neural networks, 2025
Adam Scherlis, Kshitij Sachan, Adam S. Jermyn, Joe Benton, and Buck Shlegeris · 2025
Closest in time.