Fetching the paper…
Reading the bibliography…
In this work we investigate the reasons why Batch Normalization (BN) improves the generalization performance of deep networks.
In: NIPS, pp. 2348–2356 (2011)
Graves, A.: Practical variational inference for neural networks · 2011
Earlier work this paper cites.
Lee, P.: Bayesian Statistics: An Introduction (2012)
2012
Earlier work this paper cites.
Bioinformatics 30(11), 1609–1617 (2014)
Maka, M., et al · 2014
Earlier work this paper cites.
JMLR 15, 1929–1958 (2014)
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., Salakhutdinov, R.: Dropout: A simple way to prevent neural networks from overfitting · 2014
Earlier work this paper cites.
In: ICML. pp. 1613–1622 (2015)
Blundell, C., Cornebise, J., Kavukcuoglu, K., Wierstra, D.: Weight uncertainty in neural networks · 2015
Earlier work this paper cites.
In: ICML. vol. 37, pp. 448–456 (2015)
Ioffe, S., Szegedy, C.: Batch normalization: Accelerating deep network training by reducing internal covariate shift · 2015
Earlier work this paper cites.
In: NIPS, pp. 2575–2583 (2015)
Kingma, D.P., Salimans, T., Welling, M.: Variational dropout and the local reparameterization trick · 2015
Earlier work this paper cites.
Luenberger, D.G., Ye, Y.: Linear and Nonlinear Programming (2015)
2015
Earlier work this paper cites.
In: MICCAI. pp. 234–241 (2015)
Ronneberger, O., Fischer, P., Brox, T.: U-Net: Convolutional networks for biomedical image segmentation · 2015
Cited alongside, same era.
In: NIPS, pp. 3528–3536 (2015)
Schulman, J., Heess, N., Weber, T., Abbeel, P.: Gradient estimation using stochastic computation graphs · 2015
Cited alongside, same era.
In: ICLR (workshop track) (2015)
Springenberg, J., Dosovitskiy, A., Brox, T., Riedmiller, M.: Striving for simplicity: The all convolutional net · 2015
Cited alongside, same era.
In: ICML. pp. 1168–1176 (2016)
Arpit, D., Zhou, Y., Kota, B.U., Govindaraju, V.: Normalization propagation: A parametric technique for removing internal covariate shift in deep networks · 2016
Cited alongside, same era.
In: ICLR (2016)
Clevert, D.A., Unterthiner, T., Hochreiter, S.: Fast and accurate deep network learning by exponential linear units (ELUs) · 2016
Cited alongside, same era.
In: CVPR. pp. 770–778 (2016)
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition · 2016
Cited alongside, same era.
In: BMVC. pp. 87.1–87.12 (September 2016)
Zagoruyko, S., Komodakis, N.: Wide residual networks · 2016
Later among the works it cites.
Gitman, I., Ginsburg, B.: Comparison of batch normalization and weight normalization algorithms for the large-scale image classification · 2017
Later among the works it cites.
In: ICLR Workshop track (2018)
Atanov, A., Ashukha, A., Molchanov, D., Neklyudov, K., Vetrov, D.: Uncertainty estimation via stochastic batch normalization · 2018
Closest in time.
In: CVPR (June 2018)
Gast, J., Roth, S.: Lightweight probabilistic deep networks · 2018
Closest in time.
Li, X., Chen, S., Hu, X., Yang, J.: Understanding the disharmony between dropout and batch normalization by variance shift · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ArXiv e-prints (Jul 2016)
Lei Ba, J., Kiros, J.R., Hinton, G.E.: Layer Normalization · 2016
Cited alongside, same era.
In: NIPS (2016)
Salimans, T., Kingma, D.P.: Weight normalization: A simple reparameterization to accelerate training of deep neural networks · 2016
Cited alongside, same era.
In: ICML. pp. 1050–1059 (2016a)
Gal, Y., Ghahramani, Z.: Dropout as a Bayesian approximation: Representing model uncertainty in deep learning
Cited in the paper.
In: NIPS. pp. 1027–1035 (2016b)
Gal, Y., Ghahramani, Z.: A theoretically grounded application of dropout in recurrent neural networks
Cited in the paper.
CoRR 1805.11604 (2018)
Santurkar, S., Tsipras, D., Ilyas, A., Madry, A.: How does batch normalization help optimization? (no, it is not about internal covariate shift) · 2018
Closest in time.
In: Computer Vision Winter Workshop. pp. 45–53 (2018)
Shekhovtsov, A., Flach, B.: Normalization of neural networks using analytic variance propagation · 2018
Closest in time.
In: ICML (2018)
Teye, M., Azizpour, H., Smith, K.: Bayesian uncertainty estimation for batch normalized deep networks · 2018
Closest in time.