Fetching the paper…
Reading the bibliography…
Normalization techniques are essential for accelerating the training and improving the generalization of deep neural networks (DNNs), and have successfully been used in various applications.
A. P. Dempster, N. M. Laird, and D. B. Rubin, “Maximum likelihood from incomplete data via the em algorithm,” JOURNAL OF THE ROYAL STATISTICAL SOCIETY, SERIES B , vol. 39, no. 1, pp. 1–38, 1977
1977
Earlier work this paper cites.
Y. LeCun, I. Kanter, and S. A. Solla, “Second order properties of error surfaces,” in NeurIPS , 1990
1990
Earlier work this paper cites.
A. Krogh and J. A. Hertz, “A simple weight decay can improve generalization,” in NeurIPS , 1992
1992
Earlier work this paper cites.
Y. LeCun, L. Bottou, G. B. Orr, and K.-R. Muller, “Efficient backprop,” in Neural Networks: Tricks of the Trade , 1998
1998
Earlier work this paper cites.
N. N. Schraudolph, “Accelerated gradient descent by factor-centering decomposition,” Tech. Rep., 1998
1998
Earlier work this paper cites.
G. E. Hinton and R. R. Salakhutdinov, “Reducing the dimensionality of data with neural networks,” science , vol. 313, no. 5786, pp. 504–507, 2006
2006
Earlier work this paper cites.
D. Arthur and S. Vassilvitskii, “K-means++: The advantages of careful seeding,” in Proceedings of the Eighteenth Annual ACM-SIAM Symposium on Discrete Algorithms , 2007
2007
Earlier work this paper cites.
Siwei Lyu and E. P. Simoncelli, “Nonlinear image representation using divisive normalization,” in CVPR , 2008
2008
Earlier work this paper cites.
N. J. Higham, Functions of matrices: theory and computation . SIAM, 2008
2008
Earlier work this paper cites.
A. Krizhevsky, “Learning multiple layers of features from tiny images,” Tech. Rep., 2009
2009
Earlier work this paper cites.
K. Jarrett, K. Kavukcuoglu, M. Ranzato, and Y. LeCun, “What is the best multi-stage architecture for object recognition?” in ICCV , 2009
2009
Earlier work this paper cites.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in AISTATS , 2010
2010
Earlier work this paper cites.
V. Nair and G. E. Hinton, “Rectified linear units improve restricted boltzmann machines,” in ICML , 2010
2010
Earlier work this paper cites.
J. Martens, “Deep learning via hessian-free optimization,” in ICML . Omnipress, 2010, pp. 735–742
2010
Earlier work this paper cites.
J. Duchi, E. Hazan, and Y. Singer, “Adaptive subgradient methods for online learning and stochastic optimization.” Journal of machine learning research , vol. 12, no. 7, 2011
2011
Earlier work this paper cites.
G. Montavon and K.-R. Müller, Deep Boltzmann Machines and the Centering Trick , 2012, vol. 7700
2012
Earlier work this paper cites.
T. Raiko, H. Valpola, and Y. LeCun, “Deep learning made easier by linear transformations in perceptrons,” in AISTATS , 2012
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in NeurIPS , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
O. Vinyals and D. Povey, “Krylov subspace descent for deep learning,” in AISTATS , 2012, pp. 1261–1268
2012
Earlier work this paper cites.
J. Martens and I. Sutskever, “Training deep and recurrent networks with hessian-free optimization,” in Neural Networks: Tricks of the Trade (2nd ed.) , ser. Lecture Notes in Computer Science, vol. 7700. Springer, 2012, pp. 479–535
2012
Earlier work this paper cites.
G. Hinton, N. Srivastava, and K. Swersky, “Neural networks for machine learning lecture 6a overview of mini-batch gradient descent,” 2012
2012
Earlier work this paper cites.
R. Pascanu, T. Mikolov, and Y. Bengio, “On the difficulty of training recurrent neural networks,” in ICML , 2013
2013
Earlier work this paper cites.
T. Kaneko, S. G. O. Fiori, and T. Tanaka, “Empirical arithmetic averaging over the compact stiefel manifold,” IEEE Trans. Signal Processing , vol. 61, no. 4, pp. 883–894, 2013
2013
Earlier work this paper cites.
Z. Wen and W. Yin, “A feasible method for optimization with orthogonality constraints.” Math. Program. , vol. 142, no. 1-2, pp. 397–434, 2013
2013
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in ECCV , 2014, pp. 740–755
2014
Earlier work this paper cites.
J. Martens, “New perspectives on the natural gradient method,” arXiv preprint arXiv:1412.1193 , 2014
2014
Earlier work this paper cites.
A. M. Saxe, J. L. McClelland, and S. Ganguli, “Exact solutions to the nonlinear dynamics of learning in deep linear neural networks,” in ICLR , 2014
2014
Earlier work this paper cites.
S. Wiesler, A. Richard, R. Schlüter, and H. Ney, “Mean-normalized stochastic gradient for large-scale deep learning,” in ICASSP , 2014
2014
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A simple way to prevent neural networks from overfitting,” J. Mach. Learn. Res. , vol. 15, no. 1, pp. 1929–1958, Jan. 2014
2014
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in NeurIPS , 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
G. F. Montufar, R. Pascanu, K. Cho, and Y. Bengio, “On the number of linear regions of deep neural networks,” in NeurIPS , 2014
2014
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” nature , vol. 521, no. 7553, pp. 436–444, 2015
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in ICLR , 2015
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in ICML , 2015
2015
Earlier work this paper cites.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in CVPR , 2015
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al. , “Imagenet large scale visual recognition challenge,” International journal of computer vision , vol. 115, no. 3, pp. 211–252, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
J. Martens and R. Grosse, “Optimizing neural networks with kronecker-factored approximate curvature,” in ICML , 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,” in ICCV , 2015
2015
Earlier work this paper cites.
G. Desjardins, K. Simonyan, R. Pascanu, and k. kavukcuoglu, “Natural neural networks,” in NeurIPS , 2015
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in ICLR , 2015
2015
Earlier work this paper cites.
C. Ionescu, O. Vantzos, and C. Sminchisescu, “Training deep networks with structured layers by matrix backpropagation,” in ICCV , 2015
2015
Earlier work this paper cites.
B. Neyshabur, R. Tomioka, and N. Srebro, “Norm-based capacity control in neural networks,” in COLT , 2015, pp. 1376–1401
2015
Earlier work this paper cites.
R. B. Grosse and R. Salakhutdinov, “Scaling up natural gradient by sparsely factorizing the inverse fisher matrix,” in ICML , 2015, pp. 2304–2313
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning . MIT Press, 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016
2016
Earlier work this paper cites.
S. Zagoruyko and N. Komodakis, “Wide residual networks,” in BMVC , 2016
2016
Earlier work this paper cites.
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in CVPR , 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Identity mappings in deep residual networks,” in ECCV , 2016
2016
Earlier work this paper cites.
L. J. Ba, R. Kiros, and G. E. Hinton, “Layer normalization,” arXiv preprint arXiv:1607.06450 , 2016
2016
Earlier work this paper cites.
T. Salimans and D. P. Kingma, “Weight normalization: A simple reparameterization to accelerate training of deep neural networks,” in NeurIPS , 2016
2016
Earlier work this paper cites.
D. Mishkin and J. Matas, “All you need is a good init,” in ICLR , 2016
2016
Earlier work this paper cites.
C. Laurent, G. Pereyra, P. Brakel, Y. Zhang, and Y. Bengio, “Batch normalized recurrent neural networks,” in ICASSP , 2016
2016
Earlier work this paper cites.
T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, X. Chen, and X. Chen, “Improved techniques for training gans,” in NeurIPS , 2016
2016
Earlier work this paper cites.
Ç. Gülçehre and Y. Bengio, “Knowledge matters: Importance of prior information for optimization,” The Journal of Machine Learning Research , vol. 17, no. 1, pp. 226–257, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
M. Cogswell, F. Ahmed, R. B. Girshick, L. Zitnick, and D. Batra, “Reducing overfitting in deep networks by decorrelating representations,” in ICLR , 2016
2016
Earlier work this paper cites.
W. Xiong, B. Du, L. Zhang, R. Hu, and D. Tao, “Regularizing deep convolutional neural networks with a structured decorrelation constraint,” in ICDM , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
D. Arpit, Y. Zhou, B. U. Kota, and V. Govindaraju, “Normalization propagation: A parametric technique for removing internal covariate shift in deep networks,” in ICML , 2016
2016
Earlier work this paper cites.
M. Arjovsky, A. Shah, and Y. Bengio, “Unitary evolution recurrent neural networks,” in ICML , 2016
2016
Earlier work this paper cites.
S. Wisdom, T. Powers, J. Hershey, J. Le Roux, and L. Atlas, “Full-capacity unitary recurrent neural networks,” in NeurIPS , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
B. Neyshabur, R. Tomioka, R. Salakhutdinov, and N. Srebro, “Data-dependent path normalization in neural networks,” in ICLR , 2016
2016
Earlier work this paper cites.
H. P. van Hasselt, A. Guez, M. Hessel, V. Mnih, and D. Silver, “Learning values across many orders of magnitude,” in NeurIPS , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
L. A. Gatys, A. S. Ecker, and M. Bethge, “Image style transfer using convolutional neural networks,” in CVPR , 2016
2016
Earlier work this paper cites.
S. Reed, Z. Akata, X. Yan, L. Logeswaran, B. Schiele, and H. Lee, “Generative adversarial text to image synthesis,” in ICML , 2016
2016
Earlier work this paper cites.
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” in ICLR , 2016
2016
Earlier work this paper cites.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in ICML , 2016
2016
Earlier work this paper cites.
C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “Pointnet: Deep learning on point sets for 3d classification and segmentation,” in CVPR , 2017
2017
Earlier work this paper cites.
G. Huang, Z. Liu, and K. Q. Weinberger, “Densely connected convolutional networks,” in CVPR , 2017
2017
Earlier work this paper cites.
S. Xie, R. B. Girshick, P. Dollár, Z. Tu, and K. He, “Aggregated residual transformations for deep neural networks,” in CVPR , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS , 2017
2017
Earlier work this paper cites.
J. Ba, R. Grosse, and J. Martens, “Distributed second-order optimization using kronecker-factored approximations,” in ICLR , 2017
2017
Earlier work this paper cites.
K. Sun and F. Nielsen, “Relative Fisher information and natural gradient for learning large modular models,” in ICML , 2017
2017
Earlier work this paper cites.
L. Huang, X. Liu, Y. Liu, B. Lang, and D. Tao, “Centered weight normalization in accelerating training of deep neural networks,” in ICCV , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
P. Luo, “Learning deep architectures via generalized whitened neural networks,” in ICML , 2017
2017
Earlier work this paper cites.
T. Cooijmans, N. Ballas, C. Laurent, and A. C. Courville, “Recurrent batch normalization,” in ICLR , 2017
2017
Earlier work this paper cites.
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer, “Automatic differentiation in PyTorch,” in NeurIPS Autodiff Workshop , 2017
2017
Earlier work this paper cites.
V. Dumoulin, J. Shlens, and M. Kudlur, “A learned representation for artistic style,” in ICLR , 2017
2017
Earlier work this paper cites.
X. Huang and S. Belongie, “Arbitrary style transfer in real-time with adaptive instance normalization,” in ICCV , 2017
2017
Earlier work this paper cites.
M. Ren, R. Liao, R. Urtasun, F. H. Sinz, and R. S. Zemel, “Normalizing the normalizers: Comparing and extending network normalization schemes,” in ICLR , 2017
2017
Earlier work this paper cites.
T. Kim, I. Song, and Y. Bengio, “Dynamic layer normalization for adaptive neural acoustic modeling in speech recognition,” in INTERSPEECH , 2017, pp. 2411–2415
2017
Earlier work this paper cites.
H. de Vries, F. Strub, J. Mary, H. Larochelle, O. Pietquin, and A. C. Courville, “Modulating early visual processing by language,” in NeurIPS , 2017, pp. 6594–6604
2017
Earlier work this paper cites.
B. Baker, O. Gupta, N. Naik, and R. Raskar, “Designing neural network architectures using reinforcement learning,” in ICLR , 2017
2017
Earlier work this paper cites.
S. Ioffe, “Batch renormalization: Towards reducing minibatch dependence in batch-normalized models,” in NeurIPS , 2017
2017
Earlier work this paper cites.
L. Dinh, J. Sohl-Dickstein, and S. Bengio, “Density estimation using real NVP,” in ICLR , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid scene parsing network,” in CVPR , 2017
2017
Earlier work this paper cites.
G. Klambauer, T. Unterthiner, A. Mayr, and S. Hochreiter, “Self-normalizing neural networks,” in NeurIPS , 2017
2017
Cited alongside, same era.
E. Vorontsov, C. Trabelsi, S. Kadoury, and C. Pal, “On orthogonality and learning recurrent networks with long term dependencies,” in ICML , 2017
2017
Cited alongside, same era.
S. Hyland and G. Rätsch, “Learning unitary operators with help from u(n),” in AAAI , 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
C. Moustapha, B. Piotr, E. Grave, Y. Dauphin, and N. Usunie, “Parseval networks: Improving robustness to adversarial examples,” in ICML , 2017
2017
Cited alongside, same era.
B. Ghorbani, S. Krishnan, and Y. Xiao, “An investigation into neural net optimization via hessian eigenvalue density,” in ICML , 2019
2019
Later among the works it cites.
R. Karakida, S. Akaho, and S.-i. Amari, “The normalization method for alleviating pathological sharpness in wide neural networks,” in NeurIPS , 2019, pp. 6403–6413
2019
Later among the works it cites.
2019
Later among the works it cites.
G. Yang, J. Pennington, V. Rao, J. Sohl-Dickstein, and S. S. Schoenholz, “A mean field theory of batch normalization,” in ICLR , 2019
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Brock, T. Lim, J. M. Ritchie, and N. Weston, “Neural photo editing with introspective adversarial networks,” in ICLR , 2017
2017
Cited alongside, same era.
M. Cho and J. Lee, “Riemannian approach to batch normalization,” in NeurIPS , 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
K. Jia, “Improving training of deep neural networks via singular value bounding,” in CVPR , 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
E. Hoffer, I. Hubara, and D. Soudry, “Train longer, generalize better: closing the generalization gap in large batch training of neural networks,” in NeurIPS , 2017
2017
Cited alongside, same era.
2019
Later among the works it cites.
A. Labatie, “Characterizing well-behaved vs. pathological deep neural networks,” in ICML , 2019, pp. 3611–3621
2019
Later among the works it cites.
J. Kohler, H. Daneshmand, A. Lucchi, T. Hofmann, M. Zhou, and K. Neymeyr, “Exponential convergence rates for batch normalization: The power of length-direction decoupling in non-convex optimization,” in AISTATS , 2019
2019
Later among the works it cites.
P. Luo, X. Wang, W. Shao, and Z. Peng, “Towards understanding regularization in batch normalization,” in ICLR , 2019
2019
Later among the works it cites.
X. Li, S. Chen, X. Hu, and J. Yang, “Understanding the disharmony between dropout and batch normalization by variance shift,” in CVPR , 2019
2019
Later among the works it cites.
J. Gordon, J. Bronskill, M. Bauer, S. Nowozin, and R. Turner, “Meta-learning probabilistic inference for prediction,” in ICLR , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
D. Brooks, O. Schwander, F. Barbaresco, J.-Y. Schneider, and M. Cord, “Riemannian batch normalization for spd neural networks,” in NeurIPS , 2019, pp. 15 463–15 474
2019
Later among the works it cites.
2019
Later among the works it cites.
W. Chang, T. You, S. Seo, S. Kwak, and B. Han, “Domain-specific batch normalization for unsupervised domain adaptation,” in CVPR , 2019
2019
Later among the works it cites.
R. Romijnders, P. Meletis, and G. Dubbelman, “A domain agnostic normalization layer for unsupervised adversarial domain adaptation,” in WACV , 2019
2019
Later among the works it cites.
S. Roy, A. Siarohin, E. Sangineto, S. R. Bulo, N. Sebe, and E. Ricci, “Unsupervised domain adaptation using feature-whitening and consensus loss,” in CVPR , 2019
2019
Later among the works it cites.
X. Wang, Y. Jin, M. Long, J. Wang, and M. I. Jordan, “Transferable normalization: Towards improving transferability of deep neural networks,” in NeurIPS , 2019
2019
Later among the works it cites.
Y. Li and N. Vasconcelos, “Efficient multi-domain learning by covariance normalization,” in CVPR , 2019
2019
Later among the works it cites.
Y. Jing, Y. Yang, Z. Feng, J. Ye, Y. Yu, and M. Song, “Neural style transfer: A review,” IEEE transactions on visualization and computer graphics , 2019
2019
Later among the works it cites.
T.-Y. Chiu, “Understanding generalized whitening and coloring transform for universal style transfer,” in ICCV , 2019
2019
Later among the works it cites.
W. Cho, S. Choi, D. K. Park, I. Shin, and J. Choo, “Image-to-image translation via group-wise deep whitening-and-coloring transformation,” in CVPR , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
H. Zhang, I. Goodfellow, D. Metaxas, and A. Odena, “Self-attention generative adversarial networks,” in ICML , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
T. Karras, S. Laine, and T. Aila, “A style-based generator architecture for generative adversarial networks,” in CVPR , 2019
2019
Later among the works it cites.
T. Chen, M. Lucic, N. Houlsby, and S. Gelly, “On self modulation for generative adversarial networks,” in ICLR , 2019
2019
Later among the works it cites.
J. Yu, L. Yang, N. Xu, J. Yang, and T. Huang, “Slimmable neural networks,” in ICLR , 2019
2019
Later among the works it cites.
A. Ardakani, Z. Ji, S. C. Smithson, B. H. Meyer, and W. J. Gross, “Learning recurrent binary/ternary weights,” in ICLR , 2019
2019
Later among the works it cites.
L. Hou, J. Zhu, J. Kwok, F. Gao, T. Qin, and T.-Y. Liu, “Normalization helps training of quantized lstm,” in NeurIPS , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
C. Anil, J. Lucas, and R. Grosse, “Sorting out lipschitz function approximation,” in ICLR , 2019
2019
Later among the works it cites.
H. Qian and M. N. Wegman, “L2-nonexpansive neural networks,” in ICLR , 2019
2019
Later among the works it cites.
L. Huang, L. Liu, F. Zhu, D. Wan, Z. Yuan, B. Li, and L. Shao, “Controllable orthogonalization in training dnns,” in CVPR , 2020
2020
Closest in time.
R. Xiong, Y. Yang, D. He, K. Zheng, S. Zheng, C. Xing, H. Zhang, Y. Lan, L. Wang, and T.-Y. Liu, “On layer normalization in the transformer architecture,” in ICML , 2020
2020
Closest in time.
L. Huang, L. Zhao, Y. Zhou, F. Zhu, L. Liu, and L. Shao, “An investigation into the stochasticity of batch whitening,” in CVPR , 2020
2020
Closest in time.
L. Huang, J. Qin, L. Liu, F. Zhu, and L. Shao, “Layer-wise conditioning analysis in exploring the learning dynamics of dnns,” in ECCV , 2020
2020
Closest in time.
P. A. Sokol and I. M. Park, “Information geometry of orthogonal initializations and training,” in ICLR , 2020
2020
Closest in time.
J. Bronskill, J. Gordon, J. Requeima, S. Nowozin, and R. E. Turner, “Tasknorm: Rethinking batch normalization for meta-learning,” in ICML , 2020
2020
Closest in time.
C. Summers and M. J. Dinneen, “Four things everyone should know to improve batch normalization,” in ICLR , 2020
2020
Closest in time.
C. Ye, M. Evanusa, H. He, A. Mitrokhin, T. Goldstein, J. A. Yorke, C. Fermuller, and Y. Aloimonos, “Network deconvolution,” in ICLR , 2020
2020
Closest in time.
T. Joo, D. Kang, and B. Kim, “Regularizing activations in neural networks via distribution matching with the wasserstein metric,” in ICLR , 2020
2020
Closest in time.
2020
Closest in time.
W. Shao, S. Tang, X. Pan, P. Tan, X. Wang, and P. Luo, “Channel equilibrium networks for learning deep representation,” in ICML , 2020
2020
Closest in time.
2020
Closest in time.
J. Yan, R. Wan, X. Zhang, W. Zhang, Y. Wei, and J. Sun, “Towards stabilizing batch statistics in backward propagation of batch normalization,” in ICLR , 2020
2020
Closest in time.
S. Shen, Z. Yao, A. Gholami, M. W. Mahoney, and K. Keutzer, “Powernorm: Rethinking batch normalization in transformers,” in ICML , 2020
2020
Closest in time.
2020
Closest in time.
S. Singh and S. Krishnan, “Filter response normalization layer: Eliminating batch dependence in the training of deep neural networks,” in CVPR , 2020
2020
Closest in time.
J. Kim, M. Kim, H. Kang, and K. H. Lee, “U-gat-it: Unsupervised generative attentional networks with adaptive layer-instance normalization for image-to-image translation,” in ICLR , 2020
2020
Closest in time.
S. Liang, Z. Huang, M. Liang, and H. Yang, “Instance enhancement batch normalization: an adaptive regulator of batch noise,” in AAAI , 2020
2020
Closest in time.
R. Zhang, Z. Peng, L. Wu, Z. Li, and P. Luo, “Exemplar normalization for learning deep representation,” in CVPR , 2020
2020
Closest in time.
J. Bronskill, J. Gordon, J. Requeima, S. Nowozin, and R. E. Turner, “Tasknorm: Rethinking batch normalization for meta-learning,” in ICML , 2020
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
H. Yong, J. Huang, D. Meng, X. Hua, and L. Zhang, “Momentum batch normalization for deep learning with small batch size,” in ECCV , 2020
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
X. Li, S. Chen, and J. Yang, “Understanding the disharmony between weight normalization family and weight decay.” in AAAI , 2020
2020
Closest in time.
J. Wang, Y. Chen, R. Chakraborty, and S. X. Yu, “Orthogonal convolutional neural networks,” in CVPR , 2020
2020
Closest in time.
H. Qi, C. You, X. Wang, Y. Ma, and J. Malik, “Deep isometric learning for visual recognition,” in ICML , 2020
2020
Closest in time.
J. Li, L. Fuxin, and S. Todorovic, “Efficient riemannian optimization on the stiefel manifold via the cayley transform,” in ICLR , 2020
2020
Closest in time.
Y. You, J. Li, S. Reddi, J. Hseu, S. Kumar, S. Bhojanapalli, X. Song, J. Demmel, K. Keutzer, and C.-J. Hsieh, “Large batch optimization for deep learning: Training bert in 76 minutes,” in ICLR , 2020
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
H. Yong, J. Huang, X. Hua, and L. Zhang, “Gradient centralization: A new optimization technique for deep neural networks,” in ECCV , 2020
2020
Closest in time.
J. Sun, X. Cao, H. Liang, W. Huang, Z. Chen, and Z. Li, “New interpretations of normalization methods in deep learning,” in AAAI , 2020
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
Z. Li and S. Arora, “An exponential learning rate schedule for batch normalized networks,” in ICLR , 2020
2020
Closest in time.
2020
Closest in time.
Y. Dukler, Q. Gu, and G. Montúfar, “Optimization theory for relu neural networks trained with normalization layers,” in ICML , 2020
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
C. Wang, Y. Wu, Q. Vuong, and K. Ross, “Striving for simplicity and performance in off-policy drl: Output normalization and non-uniform sampling,” in ICML , 2020
2020
Closest in time.
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in CVPR , 2020
2020
Closest in time.
M. Taha Kocyigit, L. Sevilla-Lara, T. M. Hospedales, and H. Bilen, “Unsupervised batch normalization,” in CVPR Workshops , 2020
2020
Closest in time.
W. Sun, W. Jiang, E. Trulls, A. Tagliasacchi, and K. M. Yi, “Attentive context normalization for robust permutation-equivariant learning,” in CVPR , 2020
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
C. Xie and A. Yuille, “Intriguing properties of adversarial training at scale,” in ICLR , 2020
2020
Closest in time.
C. Xie, M. Tan, B. Gong, J. Wang, A. Yuille, and Q. V. Le, “Adversarial examples improve image recognition,” in ICLR , 2020
2020
Closest in time.
S. Seo, Y. Suh, D. Kim, J. Han, and B. Han, “Learning to optimize domain specific normalization for domain generalization,” in ECCV , 2020
2020
Closest in time.
2020
Closest in time.
Y. Jing, X. Liu, Y. Ding, X. Wang, E. Ding, M. Song, and S. Wen, “Dynamic instance normalization for arbitrary style transfer,” in AAAI , 2020
2020
Closest in time.
T. Yu, Z. Guo, X. Jin, S. Wu, Z. Chen, W. Li, Z. Zhang, and S. Liu, “Region normalization for image inpainting,” in AAAI , 2020
2020
Closest in time.
Y. Wang, Y.-C. Chen, X. Zhang, J. Sun, and J. Jia, “Attentive normalization for conditional image generation,” in CVPR , 2020
2020
Closest in time.
B. Liu, Y. Zhu, Z. Fu, G. de Melo, and A. Elgammal, “Oogan: Disentangling gan with one-hot sampling and orthogonal regularization.” in AAAI , 2020
2020
Closest in time.
B. Li, B. Wu, J. Su, G. Wang, and L. Lin, “Eagleeye: Fast sub-net evaluation for efficient neural network pruning,” in ECCV , 2020
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
W. Sun, W. Jiang, E. Trulls, A. Tagliasacchi, and K. M. Yi, “Attentive context normalization for robust permutation-equivariant learning,” in CVPR , 2020
2020
Closest in time.
H. Xiong, L. Huang, M. Yu, L. Liu, F. Zhu, and L. Shao, “On the number of linear regions of convolutional neural networks,” in ICML , 2020
2020
Closest in time.