Fetching the paper…
Reading the bibliography…
A big mystery in deep learning continues to be the ability of methods to generalize when the number of model parameters is larger than the number of training examples.
A learning algorithm for boltzmann machines
D. H. Ackley, G. E. Hinton, and T. J. Sejnowski · 1985
Earlier work this paper cites.
Image compression by back propagation: An example of extensional progamming
G. Cottrell, P. Munro, and D. Zipser · 1987
Earlier work this paper cites.
Learning the hidden structure of speech
J. L. Elman and D. Zipser · 1988
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
P. Baldi and K. Hornik · 1989
Earlier work this paper cites.
Learning lie groups for invariant visual perception
R. Rao and D. L. Ruderman · 1999
Earlier work this paper cites.
A global geometric framework for nonlinear dimensionality reduction
J. B. Tenenbaum, V. De Silva, and J. C. Langford · 2000
Earlier work this paper cites.
Matching categorical object representations in inferior temporal cortex of man and monkey
N. Kriegeskorte, M. Mur, D. A. Ruff, R. Kiani, J. Bodurka, H. Esteky, K. Tanaka, and P. A. Bandettini · 2008
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
P. Vincent, H. Larochelle, Y. Bengio, and P. A. Manzagol · 2008
Earlier work this paper cites.
What is the best multi-stage architecture for object recognition?
K. Jarrett, K. Kavukcuoglu, M. A. Ranzato, and Y. LeCun · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Why does unsupervised pre-training help deep learning?
D. Erhan, Y. Bengio, A. Courville, P. A. Manzagol, P. Vincent, and S. Bengio · 2010
Earlier work this paper cites.
An unsupervised algorithm for learning lie group transformations
J. Sohl-Dickstein, C. M. Wang, and B. A. Olshausen · 2010
Earlier work this paper cites.
Lie group transformation models for predictive video coding
C. M. Wang, J. Shol-Dickstein, I. Tosic, and B. A. Olshausen · 2011
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research [best of the web]
L. Deng · 2012
Earlier work this paper cites.
Invariant scattering convolution networks
J. Bruna and S. Mallat · 2013
Earlier work this paper cites.
A. Makhzani and B. Frey · 2013
Earlier work this paper cites.
Dropout training as adaptive regularization
S. Wager, S. Wang, and P. S. Liang · 2013
Earlier work this paper cites.
Deep scattering spectrum
J. Andén and S. Mallat · 2014
Cited alongside, same era.
Unsupervised deep haar scattering on graphs
X. Chen, X. Cheng, and S. Mallat · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
On the number of linear regions of deep neural networks
G. F. Montufar, R. Pascanu, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
Why does deep learning work? a perspective from group theory
A. Paul and S. Venkatasubramanian · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Deep convolutional autoencoder-based lossy image compression
Z. Cheng, H. Sun, M. Takeuchi, and J. Katto · 2018
Later among the works it cites.
T. Cohen, M. Geiger, J. Köhler, and M. Welling · 2018
Later among the works it cites.
Explorations in homeomorphic variational auto-encoding
L. Falorsi, P. de Haan, T. R. Davidson, N. De Cao, M. Weiler, P. Forré, and T. S. Cohen · 2018
Later among the works it cites.
On the generalization of equivariance and convolution in neural networks to the action of compact groups
R. Kondor and S. Trivedi · 2018
Later among the works it cites.
On the spectral bias of neural networks
N. Rahaman, A. Baratin, D. Arpit, F. Draxler, M. Lin, F. A Hamprecht, Y. Bengio, and A. Courville · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Generalized autoencoder: A neural network framework for dimensionality reduction
W. Wang, Y. Huang, Y. Wang, and L. Wang · 2014
Cited alongside, same era.
Joint time-frequency scattering for audio classification
J. Andén, V. Lostanlen, and S. Mallat · 2015
Cited alongside, same era.
Lie Groups, Lie Algebras, and Representations: an Elementary Introduction , volume 222
B. Hall · 2015
Cited alongside, same era.
Group equivariant convolutional networks
T. Cohen and M. Welling · 2016
Cited alongside, same era.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Cited alongside, same era.
Understanding deep convolutional networks
S. Mallat · 2016
Cited alongside, same era.
Later among the works it cites.
Manifold-tiling localized receptive fields are optimal in similarity-preserving neural networks
A. Sengupta, C. Pehlevan, M. Tepper, A. Genkin, and D. Chklovskii · 2018
Later among the works it cites.
A similarity-preserving network trained on transformed images recapitulates salient features of the fly motion detection circuit
Y. Bahroun, D. Chklovskii, and A. Sengupta · 2019
Later among the works it cites.
The geometry of deep networks: power diagram subdivision
R. Balestriero, R. Cosentino, B. Aazhang, and R. Baraniuk · 2019
Later among the works it cites.
Single-cell rna-seq denoising using a deep count autoencoder
G. Eraslan, L. M. Simon, M. Mircea, N. S. Mueller, and F. J. Theis · 2019
Later among the works it cites.
On random deep weight-tied autoencoders: Exact asymptotic analysis, phase transitions, and implications to training
P. Li and P. M. Nguyen · 2019
Later among the works it cites.
On the dynamics of gradient descent for autoencoders
T. V. Nguyen, Raymond K. W. Wong, and C. Hegde · 2019
Later among the works it cites.
Learnable group transform for time-series
R. Cosentino and B. Aazhang · 2020
Closest in time.
Convex geometry of two-layer relu networks: Implicit autoencoding and interpretable models
T. Ergen and M. Pilanci · 2020
Closest in time.
Hyperplane arrangements of trained convnets are biased
M. Gamba, S. Carlsson, H. Azizpour, and M. Björkman · 2020
Closest in time.
Learning a lie algebra from unlabeled data pairs
C. Ick and V. Lostanlen · 2020
Closest in time.
A geometric understanding of deep learning
N. Lei, D. An, Y. Guo, K. Su, S. Liu, Z. Luo, S. Yau, and X. Gu · 2020
Closest in time.
Robust and interpretable blind image denoising via bias-free convolutional neural networks
S. Mohan, Z. Kadkhodaie, E. P. Simoncelli, and C. Fernandez-Granda · 2020
Closest in time.