Fetching the paper…
Reading the bibliography…
Recent work suggests that convolutional neural networks of different architectures learn to classify images in the same order.
Neural networks and principal component analysis: Learning from examples without local minima
Pierre Baldi and Kurt Hornik · 1989
Earlier work this paper cites.
Neural network ensembles, cross validation, and active learning
Anders Krogh and Jesper Vedelsby · 1994
Earlier work this paper cites.
Effect of batch learning in multilayer neural networks
Kenji Fukumizu · 1998
Earlier work this paper cites.
Natural image statistics and neural representation
Eero P Simoncelli and Bruno A Olshausen · 2001
Earlier work this paper cites.
Statistics of natural image categories
Antonio Torralba and Aude Oliva · 2003
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Self-paced learning for latent variable models
M Kumar, Benjamin Packer, and Daphne Koller · 2010
Earlier work this paper cites.
On the effectiveness of self-paced learning
Jonathan G Tullis and Aaron S Benjamin · 2011
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M. Saxe, James L. McClelland, and Surya Ganguli · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
Easy over hard: A case study on deep learning
Wei Fu and Tim Menzies · 2017
Earlier work this paper cites.
Deep nets don’t learn via memorization
David Krueger, Nicolas Ballas, Stanislaw Jastrzebski, Devansh Arpit, Maxinder S. Kanwal, Tegan Maharaj, Emmanuel Bengio, Asja Fischer, and Aaron C. Courville · 2017
Earlier work this paper cites.
Deep linear neural networks with arbitrary loss: All local minima are global
Thomas Laurent and James von Brecht · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Earlier work this paper cites.
Input–output maps are strongly biased towards simple outputs
Kamaludin Dingle, Chico Q Camargo, and Ard A Louis · 2018
Earlier work this paper cites.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Simon S. Du, Wei Hu, and Jason D. Lee · 2018
Cited alongside, same era.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason D. Lee, Daniel Soudry, and Nati Srebro · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Clément Hongler, and Franck Gabriel · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Cited alongside, same era.
Deep image prior
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky · 2018
Cited alongside, same era.
Critical points of linear neural networks: Analytical forms and landscape properties
Yi Zhou and Yingbin Liang · 2018
A mathematical theory of semantic development in deep neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2019
Later among the works it cites.
Gradient dynamics of shallow univariate relu networks
Francis Williams, Matthew Trager, Daniele Panozzo, Cláudio T. Silva, Denis Zorin, and Joan Bruna · 2019
Later among the works it cites.
On the power and limitations of random features for understanding neural networks
Gilad Yehudai and Ohad Shamir · 2019
Later among the works it cites.
Frequency bias in neural networks for input of non-uniform density
Ronen Basri, Meirav Galun, Amnon Geifman, David Jacobs, Yoni Kasten, and Shira Kritchman · 2020
Later among the works it cites.
The implicit bias of depth: How incremental learning drives generalization
Daniel Gissin, Shai Shalev-Shwartz, and Amit Daniely · 2020
Later among the works it cites.
Let’s agree to agree: Neural networks share classification order on real datasets
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2019
Cited alongside, same era.
A convergence analysis of gradient descent for deep linear neural networks
Sanjeev Arora, Nadav Cohen, Noah Golowich, and Wei Hu · 2019
Cited alongside, same era.
Learning deep linear neural networks: Riemannian gradient flows and convergence to global minimizers
Bubacarr Bah, Holger Rauhut, Ulrich Terstiege, and Michael Westdickenberg · 2019
Cited alongside, same era.
The convergence rate of neural networks for learned functions of different frequencies
Ronen Basri, David W. Jacobs, Yoni Kasten, and Shira Kritchman · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Lénaïc Chizat, Edouard Oyallon, and Francis R. Bach · 2019
Cited alongside, same era.
Width provably matters in optimization for deep linear neural networks
Simon Du and Wei Hu · 2019
Cited alongside, same era.
Guy Hacohen, Leshem Choshen, and Daphna Weinshall · 2020
Later among the works it cites.
Denoising and regularization via exploiting the structural bias of convolutional generators
Reinhard Heckel and Mahdi Soltanolkotabi · 2020
Later among the works it cites.
Provable benefit of orthogonal initialization in optimizing deep linear networks
Wei Hu, Lechao Xiao, and Jeffrey Pennington · 2020
Later among the works it cites.
Implicit bias of gradient descent for mean squared error regression with wide neural networks
Hui Jin and Guido Montúfar · 2020
Later among the works it cites.
Gradient descent with early stopping is provably robust to label noise for overparameterized neural networks
Mingchen Li, Mahdi Soltanolkotabi, and Samet Oymak · 2020
Later among the works it cites.
The pitfalls of simplicity bias in neural networks
Harshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain, and Praneeth Netrapalli · 2020
Later among the works it cites.
Theory of curriculum learning, with convex loss functions
Daphna Weinshall and Dan Amir · 2020
Later among the works it cites.
Towards understanding the spectral bias of deep learning
Yuan Cao, Zhiying Fang, Yue Wu, Ding-Xuan Zhou, and Quanquan Gu · 2021
Closest in time.
Analysis of feature learning in weight-tied autoencoders via the mean field lens
Phan-Minh Nguyen · 2021
Closest in time.
When deep classifiers agree: Analyzing correlations between learning order and image statistics
Iuliia Pliushch, Martin Mundt, Nicolas Lupp, and Visvanathan Ramesh · 2021
Closest in time.
A unifying view on implicit bias in training linear neural networks
Chulhee Yun, Shankar Krishnan, and Hossein Mobahi · 2021
Closest in time.
The grammar-learning trajectories of neural language models
Leshem Choshen, Guy Hacohen, Daphna Weinshall, and Omri Abend · 2022
Closest in time.
Active learning on a budget: Opposite strategies suit high and low budgets
Guy Hacohen, Avihu Dekel, and Daphna Weinshall · 2022
Closest in time.
Active learning through a covering lens
Ofer Yehuda, Avihu Dekel, Guy Hacohen, and Daphna Weinshall · 2022
Closest in time.