Fetching the paper…
Reading the bibliography…
We study deep neural networks (DNNs) trained on natural image data with entirely random labels.
Exponentiated gradient versus gradient descent for linear predictors
J. Kivinen and M. K. Warmuth · 1997
Earlier work this paper cites.
The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network
P. L. Bartlett · 1998
Earlier work this paper cites.
All of statistics: a concise course in statistical inference
L. Wasserman · 2004
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Robust statistics
P. J. Huber · 2011
Earlier work this paper cites.
Kernel analysis of deep networks
G. Montavon, M. L. Braun, and K.-R. Müller · 2011
Earlier work this paper cites.
Generalized cross entropy loss for training deep neural networks with noisy labels
Z. Zhang and M. Sabuncu · 2012
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
A. L. Maas, A. Y. Hannun, and A. Y. Ng · 2013
Earlier work this paper cites.
Learning with noisy labels
N. Natarajan, I. S. Dhillon, P. K. Ravikumar, and A. Tewari · 2013
Earlier work this paper cites.
Learning convolutional neural networks from few samples
R. Wagner, M. Thom, R. Schweiger, G. Palm, and A. Rothermel · 2013
Earlier work this paper cites.
Discriminative unsupervised feature learning with convolutional neural networks
A. Dosovitskiy, J. T. Springenberg, M. Riedmiller, and T. Brox · 2014
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
A. Saxe, J. McClelland, and S. Ganguli · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Learning from noisy labels with deep neural networks
S. Sukhbaatar and R. Fergus · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
J. Yosinski, J. Clune, Y. Bengio, and H. Lipson · 2014
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (ELUs)
D.-A. Clevert, T. Unterthiner, and S. Hochreiter · 2015
Earlier work this paper cites.
Fisher information distance: a geometrical reading
S. I. Costa, S. A. Santos, and J. E. Strapasson · 2015
Earlier work this paper cites.
A PCA-based convolutional network
Y. Gan, J. Liu, J. Dong, and G. Zhong · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Classification with noisy labels by importance reweighting
T. Liu and D. Tao · 2015
Cited alongside, same era.
Norm-based capacity control in neural networks
B. Neyshabur, R. Tomioka, and N. Srebro · 2015
Cited alongside, same era.
Information geometry and its applications
S. Amari · 2016
Cited alongside, same era.
Identity mappings in deep residual networks
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Co-teaching: Robust training of deep neural networks with extremely noisy labels
B. Han, Q. Yao, X. Yu, G. Niu, M. Xu, W. Hu, I. Tsang, and M. Sugiyama · 2018
Later among the works it cites.
Mentornet: Learning data-driven curriculum for very deep neural networks on corrupted labels
L. Jiang, Z. Zhou, T. Leung, L.-J. Li, and L. Fei-Fei · 2018
Later among the works it cites.
Dimensionality-driven learning with noisy labels
X. Ma, Y. Wang, M. E. Houle, S. Zhou, S. Erfani, S. Xia, S. Wijewickrema, and J. Bailey · 2018
Later among the works it cites.
Sensitivity and generalization in neural networks: an empirical study
R. Novak, Y. Bahri, D. Abolafia, J. Pennington, and J. Sohl-dickstein · 2018
Later among the works it cites.
Leveraging random label memorization for unsupervised pre-training
V. Pondenkandath, M. Alberti, S. Puran, R. Ingold, and M. Liwicki · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Convolutional neural network based on principal component analysis initialization for image classification
X.-D. Ren, H.-N. Guo, G.-C. He, X. Xu, C. Di, and S.-H. Li · 2016
Cited alongside, same era.
Wide residual networks
S. Zagoruyko and N. Komodakis · 2016
Cited alongside, same era.
A closer look at memorization in deep networks
D. Arpit, S. Jastrzębski, N. Ballas, D. Krueger, E. Bengio, M. S. Kanwal, T. Maharaj, A. Fischer, A. Courville, Y. Bengio, and S. Lacoste-Julien · 2017
Cited alongside, same era.
Information geometry
N. Ay, J. Jost, H. Vân Lê, and L. Schwachhöfer · 2017
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
P. L. Bartlett, D. J. Foster, and M. Telgarsky · 2017
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2017
Cited alongside, same era.
J. Sirignano and K. Spiliopoulos · 2018
Later among the works it cites.
Group normalization
Y. Wu and K. He · 2018
Later among the works it cites.
Critical learning periods in deep networks
A. Achille, M. Rovere, and S. Soatto · 2019
Later among the works it cites.
Intrinsic dimension of data representations in deep neural networks
A. Ansuini, A. Laio, J. H. Macke, and D. Zoccolan · 2019
Later among the works it cites.
Direct estimation of weights and efficient training of deep neural networks without sgd
N. Dehmamy, N. Rohani, and A. Katsaggelos · 2019
Later among the works it cites.
Neural network memorization dissection, 2019
J. Gu and V. Tresp · 2019
Later among the works it cites.
GPipe: Efficient training of giant neural networks using pipeline parallelism
Y. Huang, Y. Cheng, A. Bapna, O. Firat, M. X. Chen, D. Chen, H. Lee, J. Ngiam, Q. V. Le, Y. Wu, and Z. Chen · 2019
Later among the works it cites.
Big transfer (BiT): General visual representation learning
A. Kolesnikov, L. Beyer, X. Zhai, J. Puigcerver, J. Yung, S. Gelly, and N. Houlsby · 2019
Later among the works it cites.
Dying ReLU and initialization: Theory and numerical examples
L. Lu, Y. Shin, Y. Su, and G. E. Karniadakis · 2019
Later among the works it cites.
Transfusion: Understanding transfer learning for medical imaging
M. Raghu, C. Zhang, J. Kleinberg, and S. Bengio · 2019
Later among the works it cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Closest in time.
The early phase of neural network training
J. Frankle, D. J. Schwab, and A. S. Morcos · 2020
Closest in time.
Network deconvolution
C. Ye, M. Evanusa, H. He, A. Mitrokhin, T. Goldstein, J. A. Yorke, C. Fermuller, and Y. Aloimonos · 2020
Closest in time.