Fetching the paper…
Reading the bibliography…
Robust generalization is a major challenge in deep learning, particularly when the number of trainable parameters is very large.
The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network
P.L. Bartlett · 1998
Earlier work this paper cites.
Web-crawling reliability
Viv Cothey · 2004
Earlier work this paper cites.
Running experiments on amazon mechanical turk
Gabriele Paolacci, Jesse Chandler, and Panagiotis G. Ipeirotis · 2010
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors, 2012
Geoffrey E. Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan R. Salakhutdinov · 2012
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift, 2015
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Learning from massive noisy labeled data for image classification
Tong Xiao, Tian Xia, Yi Yang, Chang Huang, and Xiaogang Wang · 2015
Earlier work this paper cites.
Distinct types of eigenvector localization in networks
Romualdo Pastor-Satorras and Claudio Castellano · 2016
Earlier work this paper cites.
Robust loss functions under label noise for deep neural networks
Aritra Ghosh, Himanshu Kumar, and P Shanti Sastry · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
Masking: A new perspective of noisy supervision
Bo Han, Jiangchao Yao, Gang Niu, Mingyuan Zhou, Ivor Tsang, Ya Zhang, and Masashi Sugiyama · 2018
Earlier work this paper cites.
Not all samples are created equal: Deep learning with importance sampling
Angelos Katharopoulos and Francois Fleuret · 2018
Earlier work this paper cites.
On the importance of single directions for generalization, 2018
Ari S. Morcos, David G. T. Barrett, Neil C. Rabinowitz, and Matthew Botvinick · 2018
Earlier work this paper cites.
Deep learning from noisy image labels with quality embedding
Jiangchao Yao, Jiajie Wang, Ivor W. Tsang, Ya Zhang, Jun Sun, Chengqi Zhang, and Rui Zhang · 2018
Cited alongside, same era.
Modern Condensed Matter Physics
Steven M. Girvin and Kun Yang · 2019
Cited alongside, same era.
The generalization-stability tradeoff in neural network pruning, 2020
Brian R. Bartoldson, Ari S. Morcos, Adrian Barbu, and Gordon Erlebacher · 2020
Cited alongside, same era.
Are we done with imagenet?, 2020
Lucas Beyer, Olivier J. Hénaff, Alexander Kolesnikov, Xiaohua Zhai, and Aäron van den Oord · 2020
Cited alongside, same era.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Cited alongside, same era.
Deep learning on a data diet: Finding important examples early in training
Mansheej Paul, Surya Ganguli, and Gintare Karolina Dziugaite · 2021
Towards understanding grokking: An effective theory of representation learning
Ziming Liu, Ouail Kitouni, Niklas S Nolte, Eric Michaud, Max Tegmark, and Mike Williams · 2022
Later among the works it cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets, 2022
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Later among the works it cites.
Learning from noisy labels with deep neural networks: A survey
Hwanjun Song, Minseok Kim, Dongmin Park, Yooju Shin, and Jae-Gil Lee · 2022
Later among the works it cites.
Memorization without overfitting: Analyzing the training dynamics of large language models
Kushal Tirumala, Aram Markosyan, Luke Zettlemoyer, and Armen Aghajanyan · 2022
Later among the works it cites.
Grokking phase transitions in learning local rules with gradient descent, 2022
Bojan Žunkovič and Enej Ilievski · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Cited alongside, same era.
Hidden progress in deep learning: SGD learns parities near the computational limit
Boaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade, eran malach, and Cyril Zhang · 2022
Cited alongside, same era.
Quantifying memorization across neural language models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang · 2022
Cited alongside, same era.
Scaling laws and interpretability of learning from repeated data
Danny Hernandez, Tom Brown, Tom Conerly, Nova DasSarma, Dawn Drain, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Tom Henighan, Tristan Hume, et al · 2022
Cited alongside, same era.
Saddle-to-saddle dynamics in deep linear networks: Small initialization training, symmetry, and sparsity, 2022
Arthur Jacot, François Ged, Berfin Şimşek, Clément Hongler, and Franck Gabriel · 2022
Cited alongside, same era.
Review–a survey of learning from noisy labels
Xuefeng Liang, Xingyu Liu, and Longshan Yao · 2022
Cited alongside, same era.
Andrey Gromov · 2023
Closest in time.
Omnigrok: Grokking beyond algorithmic data
Ziming Liu, Eric J Michaud, and Max Tegmark · 2023
Closest in time.
A tale of two circuits: Grokking as competition of sparse and dense subnetworks, 2023
William Merrill, Nikolaos Tsilivis, and Aman Shukla · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability, 2023
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Closest in time.
Predicting grokking long before it happens: A look into the loss landscape of models which grok, 2023
Pascal Jr. Tikeng Notsawo, Hattie Zhou, Mohammad Pezeshki, Irina Rish, and Guillaume Dumas · 2023
Closest in time.
Explaining grokking through circuit efficiency, 2023
Vikrant Varma, Rohin Shah, Zachary Kenton, János Kramár, and Ramana Kumar · 2023
Closest in time.
The clock and the pizza: Two stories in mechanistic explanation of neural networks, 2023
Ziqian Zhong, Ziming Liu, Max Tegmark, and Jacob Andreas · 2023
Closest in time.