Fetching the paper…
Reading the bibliography…
Normalization methods are a central building block in the deep learning toolbox.
Backpropagation applied to handwritten zip code recognition
Yann LeCun, Bernhard Boser, John S Denker, Donnie Henderson, Richard E Howard, Wayne Hubbard, and Lawrence D Jackel · 1989
Earlier work this paper cites.
Adaptive mixtures of local experts
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the EM algorithm
Michael I Jordan and Robert A Jacobs · 1994
Earlier work this paper cites.
The MNIST database of handwritten digits, http://yann.lecun.com/exdb/mnist, 1998
Yann LeCun · 1998
Earlier work this paper cites.
Efficient backprop in neural networks: Tricks of the trade
Yann LeCun, Léon Bottou, Genevieve Orr, and Klaus-Robert Müller · 1998
Earlier work this paper cites.
Improving predictive inference under covariate shift by weighting the log-likelihood function
Hidetoshi Shimodaira · 2000
Earlier work this paper cites.
Mixtures of Gaussian processes
Volker Tresp · 2001
Earlier work this paper cites.
A parallel mixture of SVMs for very large scale problems
Ronan Collobert, Samy Bengio, and Yoshua Bengio · 2002
Earlier work this paper cites.
Maximum margin clustering
Linli Xu, James Neufeld, Bryce Larson, and Dale Schuurmans · 2005
Earlier work this paper cites.
Nonlinear image representation using divisive normalization
Siwei Lyu and Eero P Simoncelli · 2008
Earlier work this paper cites.
On-line expectation–maximization algorithm for latent data models
Olivier Cappé and Eric Moulines · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
What is the best multi-stage architecture for object recognition?
Kevin Jarrett, Koray Kavukcuoglu, Yann LeCun, et al · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Online EM for unsupervised models
Percy Liang and Dan Klein · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Cited alongside, same era.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Cited alongside, same era.
Learning where to attend with deep architectures for image tracking
Misha Denil, Loris Bazzani, Hugo Larochelle, and Nando de Freitas · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Cited alongside, same era.
Learning factored representations in a deep mixture of experts
David Eigen, Marc’Aurelio Ranzato, and Ilya Sutskever · 2013
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Later among the works it cites.
Universal representations: The missing link between faces, text, planktons, and cat breeds
Hakan Bilen and Andrea Vedaldi · 2017
Later among the works it cites.
Batch renormalization: Towards reducing minibatch dependence in batch-normalized models
Sergey Ioffe · 2017
Later among the works it cites.
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Later among the works it cites.
Learning multiple visual domains with residual adapters
Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi · 2017
Later among the works it cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Min Lin, Qiang Chen, and Shuicheng Yan · 2013
Cited alongside, same era.
Overfeat: Integrated recognition, localization and detection using convolutional networks
Pierre Sermanet, David Eigen, Xiang Zhang, Michaël Mathieu, Rob Fergus, and Yann LeCun · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
Conditional computation in neural networks for faster models
Emmanuel Bengio, Pierre-Luc Bacon, Joelle Pineau, and Doina Precup · 2015
Cited alongside, same era.
Natural neural networks
Guillaume Desjardins, Karen Simonyan, Razvan Pascanu, et al · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Later among the works it cites.
Mastering the game of Go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Later among the works it cites.
Improved texture networks: Maximizing quality and diversity in feed-forward stylization and texture synthesis
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky · 2017
Later among the works it cites.
Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Later among the works it cites.
Squeeze-and-excitation networks
Jie Hu, Li Shen, and Gang Sun · 2018
Closest in time.
Towards a theoretical understanding of batch normalization
Jonas Kohler, Hadi Daneshmand, Aurelien Lucchi, Ming Zhou, Klaus Neymeyr, and Thomas Hofmann · 2018
Closest in time.
Understanding the disharmony between dropout and batch normalization by variance shift
Xiang Li, Shuo Chen, Xiaolin Hu, and Jian Yang · 2018
Closest in time.
Efficient parametrization of multi-domain deep neural networks
S-A. Rebuffi, H. Bilen, and A. Vedaldi · 2018
Closest in time.
How does batch normalization help optimization? (no, it is not about internal covariate shift)
Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry · 2018
Closest in time.
Group normalization
Yuxin Wu and Kaiming He · 2018
Closest in time.