Fetching the paper…
Reading the bibliography…
Deep networks are gradually penetrating almost every domain in our lives due to their amazing success.
Robust estimation of a location parameter
Peter J Huber · 1992
Earlier work this paper cites.
On smooth activation functions
HN Mhaskar · 1997
Earlier work this paper cites.
The mnist database of handwritten digits
Yann LeCun · 1998
Earlier work this paper cites.
Beating the hold-out: Bounds for k-fold and progressive cross-validation
Avrim Blum, Adam Kalai, and John Langford · 1999
Earlier work this paper cites.
Ensemble methods in machine learning
T. G. Dietterich · 2000
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E Hinton · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Ad click prediction: a view from the trenches
H Brendan McMahan, Gary Holt, David Sculley, Michael Young, Dietmar Ebner, Julian Grady, Lan Nie, Todd Phillips, Eugene Davydov, Daniel Golovin, et al · 2013
Earlier work this paper cites.
On the number of linear regions of deep neural networks
Guido F Montufar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Deep learning with s-shaped rectified linear activation units
Xiaojie Jin, Chunyan Xu, Jiashi Feng, Yunchao Wei, Junjun Xiong, and Shuicheng Yan · 2015
Earlier work this paper cites.
Improving deep neural networks using softplus units
Hao Zheng, Zhanlei Yang, Wenju Liu, Jizhong Liang, and Yanpeng Li · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Cited alongside, same era.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Durk P Kingma · 2016
Cited alongside, same era.
Critical learning periods in deep neural networks
Alessandro Achille, Matteo Rovere, and Stefano Soatto · 2017
Cited alongside, same era.
Continuously differentiable exponential linear units
Jonathan T Barron · 2017
Cited alongside, same era.
Deep mutual learning
Ying Zhang, Tao Xiang, Timothy M Hospedales, and Huchuan Lu · 2018
Later among the works it cites.
Gradient Descent for Non-convex Problems in Modern Machine Learning
Simon Du · 2019
Later among the works it cites.
Private communication to Quoc Le
Dong Lin and Gil I. Shamir · 2019
Later among the works it cites.
Mish: A self regularized non-monotonic neural activation function
Diganta Misra · 2019
Later among the works it cites.
Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift
Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, David Sculley, Sebastian Nowozin, Joshua Dillon, Balaji Lakshminarayanan, and Jasper Snoek · 2019
Later among the works it cites.
On the invariance of the selu activation function on algorithm and hyperparameter selection in neural network recommenders
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Self-normalizing neural networks
Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter · 2017
Cited alongside, same era.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Cited alongside, same era.
Searching for activation functions
Prajit Ramachandran, Barret Zoph, and Quoc V Le · 2017
Cited alongside, same era.
An elu network with total variation for image denoising
Tianyang Wang, Zhengrui Qin, and Michelle Zhu · 2017
Cited alongside, same era.
Large scale distributed neural network training through online distillation
Rohan Anil, Gabriel Pereyra, Alexandre Passos, Robert Ormandi, George E Dahl, and Geoffrey E Hinton · 2018
Cited alongside, same era.
Deep global-connected net with the generalized multi-piecewise relu activation in deep learning
Zhi Chen and Pin-han Ho · 2018
Cited alongside, same era.
Deterministic implementations for reproducibility in deep reinforcement learning
Prabhat Nagarajan, Garrett Warnell, and Peter Stone · 2018
Cited alongside, same era.
Flora Sakketou and Nicholas Ampazis · 2019
Later among the works it cites.
A corrective view of neural networks: Representation, memorization and learning
Guy Bresler and Dheeraj Nagaraj · 2020
Closest in time.
Zhe Chen, Yuyan Wang, Dong Lin, Derek Cheng, Lichan Hong, Ed Chi, and Claire Cui · 2020
Closest in time.
Analyzing the role of model uncertainty for electronic health records
Michael W Dusenberry, Dustin Tran, Edward Choi, Jonas Kemp, Jeremy Nixon, Ghassen Jerfel, Katherine Heller, and Andrew M Dai · 2020
Closest in time.
Evolving normalization-activation layers
Hanxiao Liu, Andrew Brock, Karen Simonyan, and Quoc V Le · 2020
Closest in time.
Tanhexp: A smooth activation function with high convergence speed for lightweight neural networks
Xinyu Liu and Xiaoguang Di · 2020
Closest in time.
Generating accurate pseudo-labels in semi-supervised learning and avoiding overconfident predictions via hermite polynomial activations
Vishnu Suresh Lokhande, Songwong Tasneeyapant, Abhay Venkatesh, Sathya N Ravi, and Vikas Singh · 2020
Closest in time.
Anti-distillation: Improving reproducibility of deep networks
Gil I Shamir and Lorenzo Coviello · 2020
Closest in time.
Cihang Xie, Mingxing Tan, Boqing Gong, Alan Yuille, and Quoc V Le · 2020
Closest in time.