Fetching the paper…
Reading the bibliography…
Recent efforts at explaining the interplay of memorization and generalization in deep overparametrized networks have posited that neural networks $\textit{memorize}$ "hard" examples in the final few layers of the model.
Single units and sensation: A neuron doctrine for perceptual psychology
H. B. Barlow · 1972
Earlier work this paper cites.
Adaptive mixtures of local experts
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton · 1991
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y. Ng · 2011
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research
Li Deng · 2012
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Expert gate: Lifelong learning with a network of experts
Rahaf Aljundi, Punarjay Chakravarty, and Tinne Tuytelaars · 2017
Earlier work this paper cites.
A closer look at memorization in deep networks
Devansh Arpit, Stanisław Jastrzebski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al · 2017
Earlier work this paper cites.
Cyclical learning rates for training neural networks
Leslie N Smith · 2017
Earlier work this paper cites.
Continual learning through synaptic intelligence
Friedemann Zenke, Ben Poole, and Surya Ganguli · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Earlier work this paper cites.
Distribution density, tails, and outliers in machine learning: Metrics and applications
Nicholas Carlini, Ulfar Erlingsson, and Nicolas Papernot · 2019
Cited alongside, same era.
Superposition of many models into one
Brian Cheung, Alexander Terekhov, Yubei Chen, Pulkit Agrawal, and Bruno Olshausen · 2019
Cited alongside, same era.
An empirical study of example forgetting during deep neural network learning
Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J. Gordon · 2019
Cited alongside, same era.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2020
Cited alongside, same era.
Coherent gradients: An approach to understanding generalization in gradient descent-based optimization
Satrajit Chatterjee · 2020
Cited alongside, same era.
Extracting training data from large language models
Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Úlfar Erlingsson, Alina Oprea, and Colin Raffel · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2021
Later among the works it cites.
Characterizing structural regularities of labeled data in overparameterized models
Ziheng Jiang, Chiyuan Zhang, Kunal Talwar, and Michael C Mozer · 2021
Later among the works it cites.
On the origin of implicit regularization in stochastic gradient descent
Samuel L Smith, Benoit Dherin, David Barrett, and Soham De · 2021
Later among the works it cites.
On the geometry of generalization and memorization in deep neural networks, 2021
Cory Stephenson, Suchismita Padhy, Abhinav Ganesh, Yue Hui, Hanlin Tang, and SueYeon Chung · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Does learning require memorization? a short tale about a long tail
Vitaly Feldman · 2020
Cited alongside, same era.
What neural networks memorize and why: Discovering the long tail via influence estimation
Vitaly Feldman and Chiyuan Zhang · 2020
Cited alongside, same era.
The early phase of neural network training
Jonathan Frankle, David J. Schwab, and Ari S. Morcos · 2020
Cited alongside, same era.
Early-learning regularization prevents memorization of noisy labels
Sheng Liu, Jonathan Niles-Weed, Narges Razavian, and Carlos Fernandez-Granda · 2020
Cited alongside, same era.
The pitfalls of simplicity bias in neural networks
Harshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain, and Praneeth Netrapalli · 2020
Cited alongside, same era.
Editable neural networks
Anton Sinitsin, Vsevolod Plokhotnyuk, Dmitry Pyrkin, Sergei Popov, and Artem Babenko · 2020
Cited alongside, same era.
Deep learning through the lens of example difficulty
Robert John Nicholas Baldock, Hartmut Maennel, and Behnam Neyshabur · 2021
Cited alongside, same era.
Later among the works it cites.
Understanding invariance via feedforward inversion of discriminatively trained classifiers
Piotr Teterwak, Chiyuan Zhang, Dilip Krishnan, and Michael C Mozer · 2021
Later among the works it cites.
Modifying memories in transformer models, 2021
Chen Zhu, Ankit Singh Rawat, Manzil Zaheer, Srinadh Bhojanapalli, Daliang Li, Felix Yu, and Sanjiv Kumar · 2021
Later among the works it cites.
Reconstructing training data from trained neural networks
Niv Haim, Gal Vardi, Gilad Yehudai, Ohad Shamir, and Michal Irani · 2022
Later among the works it cites.
Robust training under label noise by over-parameterization
Sheng Liu, Zhihui Zhu, Qing Qu, and Chong You · 2022
Later among the works it cites.
Characterizing datapoints via second-split forgetting
Pratyush Maini, Saurabh Garg, Zachary Chase Lipton, and J Zico Kolter · 2022
Later among the works it cites.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Later among the works it cites.
Are all layers created equal?
Chiyuan Zhang, Samy Bengio, and Yoram Singer · 2022
Later among the works it cites.