Fetching the paper…
Reading the bibliography…
Although deep learning models have taken on commercial and political relevance, key aspects of their training and operation remain poorly understood.
On adaptive control processes
R. Bellman and R. Kalaba · 1959
Earlier work this paper cites.
A problem of dimensionality: A simple example
G. V. Trunk · 1979
Earlier work this paper cites.
Applying artificial intelligence techniques to ecological modeling
C. Loehle · 1987
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel · 1989
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Gradient-based optimization of hyperparameters
Y. Bengio · 2000
Earlier work this paper cites.
Initial sequencing and analysis of the human genome
E. S. Lander, L. M. Linton, B. Birren, C. Nusbaum, M. C. Zody, J. Baldwin, K. Devon, K. Dewar, M. Doyle, W. FitzHugh, et al · 2001
Earlier work this paper cites.
Visualizing data using t-sne
L. Van der Maaten and G. Hinton · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky, G. Hinton, et al · 2009
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2011
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Cited alongside, same era.
Gradient-based hyperparameter optimization through reversible learning
D. Maclaurin, D. Duvenaud, and R. Adams · 2015
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (elus)
D.-A. Clevert, T. Unterthiner, and S. Hochreiter · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Jamming transition as a paradigm to understand the loss landscape of deep neural networks
M. Geiger, S. Spigler, S. d’Ascoli, L. Sagun, M. Baity-Jesi, G. Biroli, and M. Wyart · 2019
Later among the works it cites.
One ticket to win them all: generalizing lottery ticket initializations across datasets and optimizers
A. Morcos, H. Yu, M. Paganini, and Y. Tian · 2019
Later among the works it cites.
Tackling climate change with machine learning
D. Rolnick, P. L. Donti, L. H. Kaack, K. Kochanski, A. Lacoste, K. Sankaran, A. S. Ross, N. Milojevic-Dupont, N. Jaques, A. Waldman-Brown, et al · 2019
Later among the works it cites.
Energy and policy considerations for deep learning in nlp
E. Strubell, A. Ganesh, and A. McCallum · 2019
Later among the works it cites.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton · 2020
Closest in time.
Deep double descent: Where bigger models and more data hurt
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms, 2017
H. Xiao, K. Rasul, and R. Vollgraf · 2017
Cited alongside, same era.
Measuring the Intrinsic Dimension of Objective Landscapes
C. Li, H. Farkhoor, R. Liu, and J. Yosinski · 2018
Cited alongside, same era.
Searching for activation functions
P. Ramachandran, B. Zoph, and Q. V. Le · 2018
Cited alongside, same era.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
M. Belkin, D. Hsu, S. Ma, and S. Mandal · 2019
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
J. Frankle and M. Carbin · 2019
Cited alongside, same era.
Hyperactivations for activation function exploration
C. J. Vercellino and W. Y. Wang
Cited in the paper.
P. Nakkiran, G. Kaplun, Y. Bansal, T. Yang, B. Barak, and I. Sutskever · 2020
Closest in time.
Synthetic petri dish: A novel surrogate model for rapid architecture search
A. Rawal, J. Lehman, F. P. Such, J. Clune, and K. O. Stanley · 2020
Closest in time.
The cost of training nlp models: A concise overview
O. Sharir, B. Peleg, and Y. Shoham · 2020
Closest in time.
A cookbook of self-supervised learning
R. Balestriero, M. Ibrahim, V. Sobal, A. Morcos, S. Shekhar, T. Goldstein, F. Bordes, A. Bardes, G. Mialon, Y. Tian, et al · 2023
Closest in time.
Guillotine regularization: Why removing layers is needed to improve generalization in self-supervised learning
F. Bordes, R. Balestriero, Q. Garrido, A. Bardes, and P. Vincent · 2023
Closest in time.
Understanding Deep Learning
S. J. Prince · 2023
Closest in time.