Fetching the paper…
Reading the bibliography…
It is widely believed that the success of deep networks lies in their ability to learn a meaningful representation of the features of the data.
Emergence of simple-cell receptive field properties by learning a sparse code for natural images
Bruno A. Olshausen and David J. Field · 1996
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Atomic decomposition by basis pursuit
Scott Shaobing Chen, David L. Donoho, and Michael A. Saunders · 1998
Earlier work this paper cites.
Regularization with dot-product kernels
Alex Smola, Zoltán Ovári, and Robert C. Williamson · 2000
Earlier work this paper cites.
Learning with kernels: support vector machines, regularization, optimization, and beyond
Bernhard Scholkopf and Alexander J Smola · 2001
Earlier work this paper cites.
Distance-based classification with lipschitz functions
Ulrike von Luxburg and Olivier Bousquet · 2004
Earlier work this paper cites.
Towards Learning Convolutions from Scratch
Behnam Neyshabur · 2007
Earlier work this paper cites.
Supervised Dictionary Learning
Julien Mairal, Jean Ponce, Guillermo Sapiro, Andrew Zisserman, and Francis Bach · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Kernel methods for deep learning
Youngmin Cho and Lawrence K. Saul · 2009
Earlier work this paper cites.
Deep Equals Shallow for ReLU Networks in Kernel Regimes
Alberto Bietti and Francis Bach · 2009
Earlier work this paper cites.
Adversarial Robustness of Supervised Sparse Coding, January 2021
Jeremias Sulam, Ramchandran Muthukumar, and Raman Arora · 2010
Earlier work this paper cites.
Spherical harmonics and approximations on the unit sphere: an introduction
Kendall Atkinson and Weimin Han · 2012
Earlier work this paper cites.
Building high-level features using large scale unsupervised learning
Quoc V Le · 2013
Earlier work this paper cites.
Invariant scattering convolution networks
Joan Bruna and Stéphane Mallat · 2013
Earlier work this paper cites.
Spherical harmonics in p dimensions
Costas Efthimiou and Christopher Frye · 2014
Earlier work this paper cites.
Norm-based capacity control in neural networks
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Earlier work this paper cites.
Understanding deep convolutional networks
Stéphane Mallat · 2016
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md Patwary, Mostofa Ali, Yang Yang, and Yanqi Zhou · 2017
Earlier work this paper cites.
Opening the black box of deep neural networks via information
Ravid Shwartz-Ziv and Naftali Tishby · 2017
Cited alongside, same era.
Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Grant M Rotskoff and Eric Vanden-Eijnden · 2018
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Implicit Bias of Gradient Descent for Wide Two-layer Neural Networks Trained with the Logistic Loss
Lénaïc Chizat and Francis Bach · 2020
Later among the works it cites.
When Do Neural Networks Outperform Kernel Methods?
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2020
Later among the works it cites.
Disentangling feature and lazy training in deep neural networks
Mario Geiger, Stefano Spigler, Arthur Jacot, and Matthieu Wyart · 2020
Later among the works it cites.
Finite versus infinite neural networks: an empirical study
Jaehoon Lee, Samuel S Schoenholz, Jeffrey Pennington, Ben Adlam, Lechao Xiao, Roman Novak, and Jascha Sohl-Dickstein · 2020
Later among the works it cites.
Scaling description of generalization with number of parameters in deep learning
Mario Geiger, Arthur Jacot, Stefano Spigler, Franck Gabriel, Levent Sagun, Stéphane d’Ascoli, Giulio Biroli, Clément Hongler, and Matthieu Wyart · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gradient descent quantizes relu network features
Hartmut Maennel, Olivier Bousquet, and Sylvain Gelly · 2018
Cited alongside, same era.
Pooling is neither necessary nor sufficient for appropriate deformation stability in CNNs
Avraham Ruderman, Neil C. Rabinowitz, Ari S. Morcos, and Daniel Zoran · 2018
Cited alongside, same era.
Intrinsic dimension of data representations in deep neural networks
Alessio Ansuini, Alessandro Laio, Jakob H Macke, and Davide Zoccolan · 2019
Cited alongside, same era.
Dimensionality compression and expansion in deep neural networks
Stefano Recanatesi, Matthew Farrell, Madhu Advani, Timothy Moore, Guillaume Lajoie, and Eric Shea-Brown · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Cited alongside, same era.
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Cited alongside, same era.
Prevalence of neural collapse during the terminal phase of deep learning training
Vardan Papyan, XY Han, and David L Donoho · 2020
Later among the works it cites.
Asymptotic learning curves of kernel methods: empirical data versus teacher–student paradigm
Stefano Spigler, Mario Geiger, and Matthieu Wyart · 2020
Later among the works it cites.
Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural Networks
Blake Bordelon, Abdulkadir Canatar, and Cengiz Pehlevan · 2020
Later among the works it cites.
Maria Refinetti, Sebastian Goldt, Florent Krzakala, and Lenka Zdeborová · 2021
Later among the works it cites.
Geometric compression of invariant manifolds in neural networks
Jonas Paccolat, Leonardo Petrini, Mario Geiger, Kevin Tyloo, and Matthieu Wyart · 2021
Later among the works it cites.
What can linearized neural networks actually say about generalization?
Guillermo Ortiz-Jiménez, Seyed-Mohsen Moosavi-Dezfooli, and Pascal Frossard · 2021
Later among the works it cites.
On the Equivalence between Neural Network and Support Vector Machine
Yilan Chen, Wei Huang, Lam M. Nguyen, and Tsui-Wei Weng · 2021
Later among the works it cites.
Landscape and training regimes in deep learning
Mario Geiger, Leonardo Petrini, and Matthieu Wyart · 2021
Later among the works it cites.
Sparse optimization on measures with over-parameterized gradient descent
Lenaic Chizat · 2021
Later among the works it cites.
Generalization error rates in kernel regression: The crossover from the noiseless to noisy regime
Hugo Cui, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová · 2021
Later among the works it cites.
Relative stability toward diffeomorphisms indicates performance in deep nets
Leonardo Petrini, Alessandro Favero, Mario Geiger, and Matthieu Wyart · 2021
Later among the works it cites.
Data-driven emergence of convolutional structure in neural networks
Alessandro Ingrosso and Sebastian Goldt · 2022
Closest in time.
Umberto M Tomasini, Antonio Sclocchi, and Matthieu Wyart · 2022
Closest in time.
Learning Theory from First Principles
Francis Bach · 2022
Closest in time.