Fetching the paper…
Reading the bibliography…
Batch Normalization (BN) has become an out-of-box technique to improve deep network training.
S. Ioffe, “Batch renormalization: Towards reducing minibatch dependence in batch-normalized models,” in Advances in Neural Information Processing Systems (NeurIPS) , 2017, pp. 1945–1953
1953
Earlier work this paper cites.
H. Wei, J. Zhang, F. Cousseau, T. Ozeki, and S.-i. Amari, “Dynamics of learning near singularities in layered networks,” Neural computation , vol. 20, no. 3, pp. 813–843, 2008
2008
Earlier work this paper cites.
A. Krizhevsky and G. Hinton, “Learning multiple layers of features from tiny images,” Master’s thesis, Department of Computer Science, University of Toronto , 2009
2009
Earlier work this paper cites.
P.-A. Absil, R. Mahony, and R. Sepulchre, Optimization algorithms on matrix manifolds . Princeton University Press, 2009
2009
Earlier work this paper cites.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics (AISTATS) , 2010, pp. 249–256
2010
Earlier work this paper cites.
V. Nair and G. E. Hinton, “Rectified linear units improve restricted boltzmann machines,” in Proceedings of the 27th International Conference on Machine Learning (ICML-10), June 21-24, 2010, Haifa, Israel , 2010. [Online]. Available: http://www.icml2010.org/papers/432.pdf
2010
Earlier work this paper cites.
B. Hariharan, P. Arbelaez, L. D. Bourdev, S. Maji, and J. Malik, “Semantic contours from inverse detectors,” in IEEE International Conference on Computer Vision (ICCV) , 2011, pp. 991–998. [Online]. Available: https://doi.org/10.1109/ICCV.2011.6126343
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems (NeurIPS) , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
Y. Nesterov, Introductory lectures on convex optimization: A basic course . Springer Science & Business Media, 2013, vol. 87
2013
Earlier work this paper cites.
T. Lin, M. Maire, S. J. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft COCO: common objects in context,” in European Conference on Computer Vision (ECCV) , 2014, pp. 740–755
2014
Earlier work this paper cites.
L. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “Semantic image segmentation with deep convolutional nets and fully connected crfs,” in International Conference on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in Proceedings of the 32nd International Conference on Machine Learning (ICML) , 2015. [Online]. Available: http://jmlr.org/proceedings/papers/v37/ioffe15.html
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei, “ImageNet Large Scale Visual Recognition Challenge,” International Journal of Computer Vision (IJCV) , vol. 115, no. 3, pp. 211–252, 2015
2015
Earlier work this paper cites.
M. Everingham, S. M. A. Eslami, L. J. V. Gool, C. K. I. Williams, J. M. Winn, and A. Zisserman, “The pascal visual object classes challenge: A retrospective,” International Journal of Computer Vision , vol. 111, no. 1, pp. 98–136, 2015. [Online]. Available: https://doi.org/10.1007/s11263-014-0733-5
2015
Earlier work this paper cites.
J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2015, pp. 3431–3440. [Online]. Available: https://doi.org/10.1109/CVPR.2015.7298965
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in International Conference on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,” in IEEE International Conference on Computer Vision (ICCV) , 2015, pp. 1026–1034
2015
Earlier work this paper cites.
S. Bubeck et al. , “Convex optimization: Algorithms and complexity,” Foundations and Trends® in Machine Learning , vol. 8, no. 3-4, pp. 231–357, 2015
2015
Earlier work this paper cites.
S. Ren, K. He, R. B. Girshick, and J. Sun, “Faster R-CNN: towards real-time object detection with region proposal networks,” in Advances in Neural Information Processing Systems (NeurIPS) , 2015, pp. 91–99. [Online]. Available: http://papers.nips.cc/paper/5638-faster-r-cnn-towards-real-time-object-detection-with-region-proposal-networks
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 770–778
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Identity mappings in deep residual networks,” European Conference on Computer Vision (ECCV) , pp. 630–645, 2016
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
S. Santurkar, D. Tsipras, A. Ilyas, and A. Madry, “How does batch normalization help optimization?” in Advances in Neural Information Processing Systems (NeurIPS) , 2018, pp. 2488–2498
2018
Later among the works it cites.
S. Qiao, C. Liu, W. Shen, and A. L. Yuille, “Few-shot image recognition by predicting parameters from activations,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018
2018
Later among the works it cites.
S. Qiao, W. Shen, Z. Zhang, B. Wang, and A. Yuille, “Deep co-training for semi-supervised image recognition,” in European Conference on Computer Vision (ECCV) , 2018
2018
Later among the works it cites.
Y. Wang, L. Xie, S. Qiao, Y. Zhang, W. Zhang, and A. L. Yuille, “Multi-scale spatially-asymmetric recalibration for image classification,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 509–525
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
D. J. Im, M. Tao, and K. Branson, “An empirical analysis of deep network loss surfaces,” 2016
2016
Cited alongside, same era.
G. Huang, Z. Liu, and K. Q. Weinberger, “Densely connected convolutional networks,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2017, pp. 2261–2269
2017
Cited alongside, same era.
R. Goyal, S. E. Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fründ, P. Yianilos, M. Mueller-Freitag, F. Hoppe, C. Thurau, I. Bax, and R. Memisevic, “The ”something something” video database for learning and evaluating visual common sense,” in IEEE International Conference on Computer Vision (ICCV) , 2017, pp. 5843–5851. [Online]. Available: https://doi.org/10.1109/ICCV.2017.622
2017
Cited alongside, same era.
W. Qiu, F. Zhong, Y. Zhang, S. Qiao, Z. Xiao, T. S. Kim, and Y. Wang, “Unrealcv: Virtual worlds for computer vision,” in Proceedings of the 25th ACM international conference on Multimedia . ACM, 2017, pp. 1221–1224
2017
Cited alongside, same era.
Y. Wang, L. Xie, C. Liu, S. Qiao, Y. Zhang, W. Zhang, Q. Tian, and A. Yuille, “SORT: Second-Order Response Transform for Visual Recognition,” IEEE International Conference on Computer Vision , 2017
2017
Cited alongside, same era.
L. Huang, X. Liu, Y. Liu, B. Lang, and D. Tao, “Centered weight normalization in accelerating training of deep neural networks,” in Proceedings of the IEEE International Conference on Computer Vision , 2017
2017
Cited alongside, same era.
M. Cho and J. Lee, “Riemannian approach to batch normalization,” in Advances in Neural Information Processing Systems , 2017, pp. 5225–5235
2017
Cited alongside, same era.
C. Yang, L. Xie, S. Qiao, and A. Yuille, “Knowledge distillation in generations: More tolerant teachers educate better students,” AAAI , 2018
2018
Later among the works it cites.
Z. Zhang, S. Qiao, C. Xie, W. Shen, B. Wang, and A. L. Yuille, “Single-shot object detection with enriched semantics,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018, pp. 5813–5821
2018
Later among the works it cites.
A. E. Orhan and X. Pitkow, “Skip connections eliminate singularities,” International Conference on Learning Representations (ICLR) , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
C. Liu, B. Zoph, M. Neumann, J. Shlens, W. Hua, L. Li, L. Fei-Fei, A. L. Yuille, J. Huang, and K. Murphy, “Progressive neural architecture search,” in European Conference on Computer Vision (ECCV) , 2018, pp. 19–35. [Online]. Available: https://doi.org/10.1007/978-3-030-01246-5_2
2018
Later among the works it cites.
C. Peng, T. Xiao, Z. Li, Y. Jiang, X. Zhang, K. Jia, G. Yu, and J. Sun, “Megdet: A large mini-batch object detector,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018, pp. 6181–6189
2018
Later among the works it cites.
2018
Later among the works it cites.
B. Zhou, A. Andonian, A. Oliva, and A. Torralba, “Temporal relational reasoning in videos,” in European Conference on Computer Vision (ECCV) , 2018, pp. 803–818
2018
Later among the works it cites.
P. Luo, X. Wang, W. Shao, and Z. Peng, “Towards understanding regularization in batch normalization,” in International Conference on Learning Representations (ICLR) , 2019
2019
Closest in time.
G. Yang, J. Pennington, V. Rao, J. Sohl-Dickstein, and S. S. Schoenholz, “A mean field theory of batch normalization,” in International Conference on Learning Representations (ICLR) , 2019
2019
Closest in time.
S. Arora, Z. Li, and K. Lyu, “Theoretical analysis of auto rate-tuning by batch normalization,” in International Conference on Learning Representations (ICLR) , 2019
2019
Closest in time.
2019
Closest in time.
P. Luo, P. Zhanglin, S. Wenqi, Z. Ruimao, R. Jiamin, and W. Lingyun, “Differentiable dynamic normalization for learning deep representation,” in International Conference on Machine Learning , 2019, pp. 4203–4211
2019
Closest in time.