Fetching the paper…
Reading the bibliography…
Normalization techniques have become a basic component in modern convolutional neural networks (ConvNets).
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,”
1998
Earlier work this paper cites.
New York, NY, USA: Springer, second ed., 2006
J. Nocedal and S. J. Wright, · 2006
Earlier work this paper cites.
A. Krizhevsky
2009
Earlier work this paper cites.
K. Gregor and Y. LeCun, “Learning fast approximations of sparse coding.,” in
2010
Earlier work this paper cites.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in
2010
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,”
2014
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,”
2014
Earlier work this paper cites.
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. C. Courville, and Y. Bengio, “Generative adversarial nets,” in
2014
Earlier work this paper cites.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein,
2015
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,”
2015
Earlier work this paper cites.
I. Goodfellow, J. Shlens, and C. Szegedy, “Explaining and harnessing adversarial examples,” in
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,”
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,”
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
M. Arjovsky, A. Shah, and Y. Bengio, “Unitary evolution recurrent neural networks,” in
2016
Earlier work this paper cites.
M. Harandi and B. Fernando, “Generalized backpropagation, Étude de cas: Orthogonality,”
2016
Earlier work this paper cites.
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He, “Aggregated residual transformations for deep neural networks,” in
2017
Earlier work this paper cites.
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in
2017
Cited alongside, same era.
E. Vorontsov, C. Trabelsi, S. Kadoury, and C. Pal, “On orthogonality and learning recurrent networks with long term dependencies,” in
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
K. Janocha and W. M. Czarnecki, “On loss functions for deep neural networks in classification,”
2017
K. Jia, S. Li, Y. Wen, T. Liu, and D. Tao, “Orthogonal deep neural networks,”
2019
Later among the works it cites.
G. Zhang, K. Niwa, and W. B. Kleijn, “Approximated orthonormal normalisation in training neural networks,” 2019
2019
Later among the works it cites.
Q. Li, S. Haque, C. Anil, J. Lucas, R. Grosse, and J.-H. Jacobsen, “Preventing gradient attenuation in lipschitz constrained convolutional networks,”
2019
Later among the works it cites.
J. Wang, Y. Chen, R. Chakraborty, and S. X. Yu, “Orthogonal convolutional neural networks,” 2019
2019
Later among the works it cites.
Q. Qu, X. Li, and Z. Zhu, “A nonconvex approach for exact and efficient multichannel sparse blind deconvolution,” in
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
D. Arpit, S. Jastrzębski, N. Ballas, D. Krueger, E. Bengio, M. S. Kanwal, T. Maharaj, A. Fischer, A. Courville, Y. Bengio,
2017
Cited alongside, same era.
G. Patrini, A. Rozza, A. Krishna Menon, R. Nock, and L. Qu, “Making deep neural networks robust to label noise: A loss correction approach,” in
2017
Cited alongside, same era.
I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. C. Courville, “Improved training of wasserstein GANs,” in
2017
Cited alongside, same era.
N. Kodali, J. Abernethy, J. Hays, and Z. Kira, “On convergence and stability of GANs,”
2017
Cited alongside, same era.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “GANs trained by a two time-scale update rule converge to a local nash equilibrium,” in
2017
Cited alongside, same era.
C. Kong and S. Lucey, “Take it in your stride: Do we need striding in cnns?,”
2017
Cited alongside, same era.
Y. Wu and K. He, “Group normalization,” in
2018
Cited alongside, same era.
2019
Later among the works it cites.
C. Guo, J. Gardner, Y. You, A. G. Wilson, and K. Weinberger, “Simple black-box adversarial attacks,” in
2019
Later among the works it cites.
A. Shafahi, M. Najibi, M. A. Ghiasi, Z. Xu, J. Dickerson, C. Studer, L. S. Davis, G. Taylor, and T. Goldstein, “Adversarial training for free!,” in
2019
Later among the works it cites.
2019
Later among the works it cites.
Z. Zhou, J. Liang, Y. Song, L. Yu, H. Wang, W. Zhang, Y. Yu, and Z. Zhang, “Lipschitz generative adversarial nets,” in
2019
Later among the works it cites.
2020
Later among the works it cites.
Y. Sun, X. Wang, Z. Liu, J. Miller, A. Efros, and M. Hardt, “Test-time training with self-supervision for generalization under distribution shifts,” in
2020
Later among the works it cites.
W. Hu, L. Xiao, and J. Pennington, “Provable benefit of orthogonal initialization in optimizing deep linear networks,” in
2020
Later among the works it cites.
H. Qi, C. You, X. Wang, Y. Ma, and J. Malik, “Deep isometric learning for visual recognition,” in
2020
Later among the works it cites.
B. Liu, Y. Zhu, Z. Fu, G. de Melo, and A. Elgammal, “Oogan: Disentangling gan with one-hot sampling and orthogonal regularization.,” in
2020
Later among the works it cites.
C. Ye, M. Evanusa, H. He, A. Mitrokhin, T. Goldstein, J. A. Yorke, C. Fermuller, and Y. Aloimonos, “Network deconvolution,” in
2020
Later among the works it cites.
M. Atzmon, A. Gropp, and Y. Lipman, “Isometric autoencoders,”
2020
Later among the works it cites.
J. Li, F. Li, and S. Todorovic, “Efficient riemannian optimization on the stiefel manifold via the cayley transform,” in
2020
Later among the works it cites.
L. Huang, L. Liu, F. Zhu, D. Wan, Z. Yuan, B. Li, and L. Shao, “Controllable orthogonalization in training dnns,” 2020
2020
Later among the works it cites.
Q. Qu, Y. Zhai, X. Li, Y. Zhang, and Z. Zhu, “Geometric analysis of nonconvex optimization landscapes for overcomplete learning,” in
2020
Later among the works it cites.
X. Chen and K. He, “Exploring simple siamese representation learning,”
2020
Later among the works it cites.
M. Li, M. Soltanolkotabi, and S. Oymak, “Gradient descent with early stopping is provably robust to label noise for overparameterized neural networks,” in
2020
Later among the works it cites.
S. Liu, J. Niles-Weed, N. Razavian, and C. Fernandez-Granda, “Early-learning regularization prevents memorization of noisy labels,”
2020
Later among the works it cites.
A. Araujo, B. Negrevergne, Y. Chevaleyre, and J. Atif, “On lipschitz regularization of convolutional layers using toeplitz matrix theory,” 2021
2021
Closest in time.
A. Trockman and J. Z. Kolter, “Orthogonalizing convolutional layers with the cayley transform,” in
2021
Closest in time.