Fetching the paper…
Reading the bibliography…
Deep learning algorithms demonstrate a surprising ability to learn high-dimensional tasks from limited examples.
D. C. Van Essen and J. H. Maunsell, Hierarchical organization and functional streams in the visual cortex, Trends in neurosciences 6
1983
Earlier work this paper cites.
E. Gardner and B. Derrida, Three unfinished works on the optimal storage capacity of networks, Journal of Physics A: Mathematical and General 22
1989
Earlier work this paper cites.
U. Grenander, Elements of pattern theory (JHU Press, 1996)
1996
Earlier work this paper cites.
G. Rozenberg and A. Salomaa, Handbook of Formal Languages (Springer, 1997)
1997
Earlier work this paper cites.
L. Györfi, M. Kohler, A. Krzyzak, H. Walk, et al. , A distribution-free theory of nonparametric regression , Vol. 1 (Springer New York, NY, 2002)
2002
Earlier work this paper cites.
U. v. Luxburg and O. Bousquet, Distance-based classification with lipschitz functions, The Journal of Machine Learning Research 5
2004
Earlier work this paper cites.
K. Grill-Spector and R. Malach, The human visual cortex, Annu. Rev. Neurosci. 27
2004
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, Imagenet: A large-scale hierarchical image database, in 2009 IEEE conference on computer vision and pattern recognition (IEEE, 2009) pp. 248–255
2009
Earlier work this paper cites.
A. Krizhevsky, Learning multiple layers of features from tiny images, Preprint at https://www.cs.toronto.edu/ kriz/learning-features-2009-TR.pdf (2009)
2009
Earlier work this paper cites.
S. Kpotufe, k-nn regression adapts to local intrinsic dimension, in Advances in Neural Information Processing Systems , Vol. 24, edited by J. Shawe-Taylor, R. Zemel, P. Bartlett, F. Pereira, and K. Weinberger (Curran Associates, Inc., 2011) pp. 729–737
2011
Earlier work this paper cites.
N. Kruger, P. Janssen, S. Kalkan, M. Lappe, A. Leonardis, J. Piater, A. J. Rodriguez-Sanchez, and L. Wiskott, Deep hierarchies in the primate visual cortex: What can we learn for computer vision?, IEEE transactions on pattern analysis and machine intelligence 35
2012
Earlier work this paper cites.
J. Bruna and S. Mallat, Invariant scattering convolution networks, IEEE transactions on pattern analysis and machine intelligence 35
2013
Earlier work this paper cites.
M. Denil, B. Shakibi, L. Dinh, M. A. Ranzato, and N. de Freitas, Predicting parameters in deep learning, in Advances in Neural Information Processing Systems , Vol. 26, edited by C. Burges, L. Bottou, M. Welling, Z. Ghahramani, and K. Weinberger (Curran Associates, Inc., 2013) pp. 2148–2156
2013
Earlier work this paper cites.
M. D. Zeiler and R. Fergus, Visualizing and understanding convolutional networks, in Computer Vision – ECCV 2014 , Lecture Notes in Computer Science (2014) pp. 818–833
2014
Earlier work this paper cites.
E. L. Denton, W. Zaremba, J. Bruna, Y. LeCun, and R. Fergus, Exploiting linear structure within convolutional networks for efficient evaluation, in Advances in Neural Information Processing Systems , Vol. 27, edited by Z. Ghahramani, M. Welling, C. Cortes, N. Lawrence, and K. Weinberger (Curran Associates, Inc., 2014) pp. 1269–1277
2014
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, Deep learning, Nature 521
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
L. Zdeborová and F. Krzakala, Statistical physics of inference: Thresholds and algorithms, Advances in Physics 65
2016
Earlier work this paper cites.
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al. , Mastering the game of go without human knowledge, Nature 550
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
H. Mhaskar, Q. Liao, and T. Poggio, When and why are deep networks better than shallow ones?, Proceedings of the AAAI Conference on Artificial Intelligence 31
2017
Cited alongside, same era.
T. Poggio, H. Mhaskar, L. Rosasco, B. Miranda, and Q. Liao, Why and when can deep-but not shallow-networks avoid the curse of dimensionality: a review, International Journal of Automation and Computing 14
2017
Cited alongside, same era.
M. Mézard, Mean-field message-passing equations in the hopfield model and its generalizations, Physical Review E 95
2017
Cited alongside, same era.
F. Bach, Breaking the curse of dimensionality with convex neural networks, Journal of Machine Learning Research 18
2017
Cited alongside, same era.
X. Yu, T. Liu, X. Wang, and D. Tao, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2017) pp. 7370–7379
S. Goldt, M. Mézard, F. Krzakala, and L. Zdeborová, Modeling the Influence of Data Structure on Learning in Neural Networks: The Hidden Manifold Model, Physical Review X 10
2020
Later among the works it cites.
2020
Later among the works it cites.
E. Malach and S. Shalev-Shwartz, The implications of local correlation on learning some deep functions, in Advances in Neural Information Processing Systems , Vol. 33, edited by H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin (Curran Associates, Inc., 2020) pp. 1322–1332
2020
Later among the works it cites.
P. Pope, C. Zhu, A. Abdelkader, M. Goldblum, and T. Goldstein, The intrinsic dimension of images and its impact on learning, in International Conference on Learning Representations (2021)
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
S. Shalev-Shwartz, O. Shamir, and S. Shammah, Failures of gradient-based deep learning, in Proceedings of the 34th International Conference on Machine Learning , Proceedings of Machine Learning Research, Vol. 70, edited by D. Precup and Y. W. Teh (PMLR, 2017) pp. 3067–3075
2017
Cited alongside, same era.
A. Voulodimos, N. Doulamis, A. Doulamis, and E. Protopapadakis, Deep learning for computer vision: A brief review, Computational Intelligence and Neuroscience , 1–13 (2018)
2018
Cited alongside, same era.
2018
Cited alongside, same era.
A. Jacot, F. Gabriel, and C. Hongler, Neural tangent kernel: Convergence and generalization in neural networks, in Advances in Neural Information Processing Systems , Vol. 31, edited by S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Curran Associates, Inc., 2018) pp. 8571–8580
2018
Cited alongside, same era.
A. M. Saxe, Y. Bansal, J. Dapello, M. Advani, A. Kolchinsky, B. D. Tracey, and D. D. Cox, On the information bottleneck theory of deep learning, Journal of Statistical Mechanics: Theory and Experiment 2019
2019
Cited alongside, same era.
A. Ansuini, A. Laio, J. H. Macke, and D. Zoccolan, Intrinsic dimension of data representations in deep neural networks, Advances in Neural Information Processing Systems 32
2019
Cited alongside, same era.
2019
Cited alongside, same era.
L. Petrini, A. Favero, M. Geiger, and M. Wyart, Relative stability toward diffeomorphisms indicates performance in deep nets, Advances in Neural Information Processing Systems 34
2021
Later among the works it cites.
F. Bach, The quest for adaptivity, Machine Learning Research Blog (2021)
2021
Later among the works it cites.
T. Hamm and I. Steinwart, Adaptive learning rates for support vector machines working on data with low intrinsic dimension, The Annals of Statistics 49
2021
Later among the works it cites.
M. Geiger, L. Petrini, and M. Wyart, Landscape and training regimes in deep learning, Physics Reports 924
2021
Later among the works it cites.
J. Paccolat, L. Petrini, M. Geiger, K. Tyloo, and M. Wyart, Geometric compression of invariant manifolds in neural networks, Journal of Statistical Mechanics: Theory and Experiment 2021
2021
Later among the works it cites.
E. Abbe, E. Boix-Adsera, M. S. Brennan, G. Bresler, and D. Nagaraj, The staircase property: How hierarchical structure can guide deep learning, in Advances in Neural Information Processing Systems , Vol. 34, edited by M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan (Curran Associates, Inc., 2021) pp. 26989–27002
2021
Later among the works it cites.
A. Favero, F. Cagnetta, and M. Wyart, Locality defeats the curse of dimensionality in convolutional teacher-student scenarios, in Advances in Neural Information Processing Systems , Vol. 34, edited by M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan (Curran Associates, Inc., 2021) pp. 9456–9467
2021
Later among the works it cites.
A. Ingrosso and S. Goldt, Data-driven emergence of convolutional structure in neural networks, Proceedings of the National Academy of Sciences 119
2022
Later among the works it cites.
B. Barak, B. Edelman, S. Goel, S. Kakade, E. Malach, and C. Zhang, Hidden progress in deep learning: Sgd learns parities near the computational limit, in Advances in Neural Information Processing Systems , Vol. 35, edited by S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Curran Associates, Inc., 2022) pp. 21750–21764
2022
Later among the works it cites.
A. Damian, J. Lee, and M. Soltanolkotabi, Neural networks can learn representations with gradient descent, in Proceedings of Thirty Fifth Conference on Learning Theory , Proceedings of Machine Learning Research, Vol. 178, edited by P.-L. Loh and M. Raginsky (PMLR, 2022) pp. 5413–5452
2022
Later among the works it cites.
J. Ba, M. A. Erdogdu, T. Suzuki, Z. Wang, D. Wu, and G. Yang, High-dimensional asymptotics of feature learning: How one gradient step improves the representation, in Advances in Neural Information Processing Systems , Vol. 35, edited by S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Curran Associates, Inc., 2022) pp. 37932–37946
2022
Later among the works it cites.
U. M. Tomasini, L. Petrini, F. Cagnetta, and M. Wyart, How deep convolutional neural networks lose spatial information with training, Machine Learning: Science and Technology 4
2023
Closest in time.
F. Cagnetta, A. Favero, and M. Wyart, What can be learnt with wide convolutional neural networks?, in Proceedings of the 40th International Conference on Machine Learning , Proceedings of Machine Learning Research, Vol. 202, edited by A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett (PMLR, 2023) pp. 3347–3379
2023
Closest in time.
2023
Closest in time.
M. Mézard, Spin glass theory and its new challenge: structured disorder, Indian Journal of Physics , 1 (2023)
2023
Closest in time.
2023
Closest in time.