Normalization propagation: A parametric technique for removing internal covariate shift in deep networks
Arpit, D., Zhou, Y., Kota, B., and Govindaraju, V · 2016
Cited alongside, same era.
Layer normalization
Original
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Cited alongside, same era.
Recurrent batch normalization
Original
Cooijmans, T., Ballas, N., Laurent, C., Gülçehre, Ç., and Courville, A · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Binarized neural networks
Hubara, I., Courbariaux, M., Soudry, D., El-Yaniv, R., and Bengio, Y · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, T. and Kingma, D. P · 2016
Cited alongside, same era.
Improved techniques for training gans
Salimans, T., Goodfellow, I., Zaremba, W., et al · 2016
Cited alongside, same era.
Instance normalization: The missing ingredient for fast stylization
Original
Ulyanov, D., Vedaldi, A., and Lempitsky, V. S · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Original
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Cited alongside, same era.
Shifting mean activation towards zero with bipolar activation functions
Original
Eidnes, L. and Nøkland, A · 2017
Cited alongside, same era.
Comparison of batch normalization and weight normalization algorithms for the large-scale image classification
Original
Gitman, I. and Ginsburg, B · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Hoffer, E., Hubara, I., and Soudry, D · 2017
Cited alongside, same era.