Fetching the paper…
Reading the bibliography…
Identifying computational mechanisms for memorization and retrieval of data is a long-standing problem at the intersection of machine learning and neuroscience.
Identity crisis: Memorization and generalization under extreme overparameterization, 2019
Chiyuan Zhang, Samy Bengio, Moritz Hardt, and Yoram Singer · 1902
Earlier work this paper cites.
Implicit generation and generalization in energy-based models, 2019
Yilun Du and Igor Mordatch · 1903
Earlier work this paper cites.
The existence of persistent states in the brain
W.A. Little · 1974
Earlier work this paper cites.
Neural Networks and Physical Systems with Emergent Collective Computational Abilities
John J. Hopfield · 1982
Earlier work this paper cites.
’Unlearning’ has a stabilizing effect in collective memories
John J. Hopfield, David I. Feinstein, and Richard G. Palmer · 1983
Earlier work this paper cites.
Competitive learning: From interactive activation to adaptive resonance
Stefan Grossberg · 1987
Earlier work this paper cites.
The Capacity of the Hopfield Associative Memory
Robert McEliece, Edward Posner, Eugene R. Rodemich, and Santosh S. Venkatesh · 1987
Earlier work this paper cites.
Bidirectional associative memories
Bart Kosko · 1988
Earlier work this paper cites.
Neural Computers
Joachim Buhmann and Klaus Schulten · 1989
Earlier work this paper cites.
A Fast Fixed-Point Algorithm for Independent Component Analysis
Aapo Hyvärinen and Erkki Oja · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
A fast learning algorithm for deep beliefnets
Geoffrey E. Hinton, Simon Osindero, and Yee-Whye Teh · 2006
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Efficient learning of deep Boltzmann machines
Ruslan Salakhutdinov and Hugo Larochelle · 2010
Cited alongside, same era.
Contractive auto-encoders: Explicit invariance during feature extraction
Salah Rifai, Pascal Vincent, Xavier Muller, Xavier Glorot, and Yoshua Bengio · 2011
Cited alongside, same era.
Autoencoders, unsupervised learning, and deep architectures
Pierre Baldi · 2012
Cited alongside, same era.
A practical guide to training restricted Boltzmann machines
Geoffrey E. Hinton · 2012
Cited alongside, same era.
What regularized auto-encoders learn from the data-generating distribution
Guillaume Alain and Yoshua Bengio · 2014
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Later among the works it cites.
Searching for activation functions
Prajit Ramachandran, Barret Zoph, and Quoc V. Le · 2017
Later among the works it cites.
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Later among the works it cites.
On the Optimization of Deep Networks: Implicit Acceleration by Overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Later among the works it cites.
Eigenvectors of Orthogonally Decomposable Functions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
U-Net: Convolutional Networks for Biomedical Image Segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Cited alongside, same era.
Nonlinear Dynamics and Chaos
Steven Strogatz · 2015
Cited alongside, same era.
Empirical evaluation of rectified activations in convolution network, 2015
Bing Xu, Naiyan Wang, Tianqi Chen, and Mu Li · 2015
Cited alongside, same era.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
Pixel recurrent neural networks
Aaron van den Oord, Nal Kalchbrenner, and Koray Kavukcuglu · 2016
Cited alongside, same era.
Mikhail Belkin, Luis Rademacher, and James Voss · 2018
Later among the works it cites.
Is it time to Swish? Comparing deep learning activation functions across NLP tasks
Steffen Eger, Paul Youssef, and Iryna Gurevych · 2018
Later among the works it cites.
Neural Tangent Kernel: Convergence and Generalization in Neural Networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Later among the works it cites.
A convergence analysis of gradient descent for deep linear neural networks
Sanjeev Arora, Nadav Cohen, Noah Golowich, and Wei Hu · 2019
Closest in time.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Closest in time.
Gradient descent finds global minima of deep neural networks
Simon S. Du, Jason D. Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Closest in time.
Gradient descent provably optimizes over-parameterized neural networks
Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Closest in time.
Global convergence of adaptive gradient methods for an over-parameterized neural network?
Xiaoxia Wu, Simon S. Du, and Rachel Ward · 2019
Closest in time.
Meta-learning deep energy-based memory models
Sergey Bartunov, Jack Rae, Simon Osindero, and Timothy Lillicrap · 2020
Closest in time.