Fetching the paper…
Reading the bibliography…
The softmax function combined with a cross-entropy loss is a principled approach to modeling probability distributions that has become ubiquitous in deep learning.
Neural Tangents: Fast and Easy Infinite Neural Networks in Python
Roman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee, Alexander A. Alemi, Jascha Sohl-Dickstein, and Samuel S. Schoenholz · 1912
Earlier work this paper cites.
Backpropagation Applied to Handwritten Zip Code Recognition
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel · 1989
Earlier work this paper cites.
Probabilistic Outputs for Support Vector Machines and Comparisons to Regularized Likelihood Methods
John Platt · 2000
Earlier work this paper cites.
Revisiting squared-error and cross-entropy functions for training neural network classifiers
Douglas M. Kline and Victor L. Berardi · 2005
Earlier work this paper cites.
Learning Multiple Layers of Features from Tiny Images
A. Krizhevsky · 2009
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Cross-Entropy vs. Squared Error Training: A Theoretical and Experimental Comparison
Pavel Golik, Patrick Doetsch, and Hermann Ney · 2013
Earlier work this paper cites.
Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Rethinking the Inception Architecture for Computer Vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Cited alongside, same era.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Sergey Zagoruyko and Nikos Komodakis · 2017
Cited alongside, same era.
Regularizing Neural Networks by Penalizing Confident Output Distributions
Gabriel Pereyra, George Tucker, Jan Chorowski, Lukasz Kaiser, and Geoffrey Hinton · 2017
Cited alongside, same era.
Classification is a Strong Baseline for Deep Metric Learning
Andrew Zhai and Hao-Yu Wu · 2019
Later among the works it cites.
Gradient Descent Provably Optimizes Over-parameterized Neural Networks
Simon S. Du, Xiyu Zhai, Barnabás Póczos, and Aarti Singh · 2019
Later among the works it cites.
Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient Descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Later among the works it cites.
Disentangling trainability and generalization in deep learning
Lechao Xiao, Jeffrey Pennington, and Samuel S. Schoenholz · 2019
Later among the works it cites.
Why bigger is not always better: On finite and infinite neural networks
Laurence Aitchison · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger · 2017
Cited alongside, same era.
Improving Generalization via Scalable Neighborhood Component Analysis
Zhirong Wu, Alexei A. Efros, and Stella X. Yu · 2018
Cited alongside, same era.
Neural Tangent Kernel: Convergence and Generalization in Neural Networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Cited alongside, same era.
DropBlock: A regularization method for convolutional networks
Golnaz Ghiasi, Tsung-Yi Lin, and Quoc V Le · 2018
Cited alongside, same era.
JAX: Composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, and Skye Wanderman-Milne · 2018
Cited alongside, same era.
When does label smoothing help?
Rafael Müller, Simon Kornblith, and Geoffrey E Hinton · 2019
Cited alongside, same era.
Bayesian Deep Convolutional Networks with Many Channels are Gaussian Processes
Roman Novak, Lechao Xiao, Jaehoon Lee, Yasaman Bahri, Greg Yang, Jiri Hron, Daniel A. Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein
Cited in the paper.
Later among the works it cites.
On Lazy Training in Differentiable Programming
Lénaïc Chizat, Edouard Oyallon, and Francis Bach · 2019
Later among the works it cites.
Reverse engineering recurrent networks for sentiment classification reveals line attractor dynamics
Niru Maheswaranathan, Alex Williams, Matthew Golub, Surya Ganguli, and David Sussillo · 2019
Later among the works it cites.
Implicit Bias of Gradient Descent for Wide Two-layer Neural Networks Trained with the Logistic Loss
Lenaic Chizat and Francis Bach · 2020
Closest in time.
Finite Versus Infinite Neural Networks: An Empirical Study
Jaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam, Lechao Xiao, Roman Novak, and Jascha Sohl-Dickstein · 2020
Closest in time.
The large learning rate phase of deep learning: The catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Closest in time.