Fetching the paper…
Reading the bibliography…
Normalization is known to help the optimization of deep neural networks.
A survey of preconditioned iterative methods for linear systems of algebraic equations
Axelsson, O · 1985
Earlier work this paper cites.
Applied numerical linear algebra , volume 56
Demmel, J. W · 1997
Earlier work this paper cites.
A new model for learning in graph domains
Gori, M., Monfardini, G., and Scarselli, F · 2005
Earlier work this paper cites.
Towards deeper graph neural networks with differentiable group normalization
Zhou, K., Huang, X., Li, Y., Zha, D., Chen, R., and Hu, X · 2006
Earlier work this paper cites.
The graph neural network model
Scarselli, F., Gori, M., Tsoi, A. C., Hagenbuchner, M., and Monfardini, G · 2008
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Weisfeiler-lehman graph kernels
Shervashidze, N., Schweitzer, P., Leeuwen, E. J. v., Mehlhorn, K., and Borgwardt, K. M · 2011
Earlier work this paper cites.
Matrix analysis
Horn, R. A. and Johnson, C. R · 2012
Earlier work this paper cites.
Spectral networks and locally connected networks on graphs
Bruna, J., Zaremba, W., Szlam, A., and LeCun, Y · 2013
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Deep graph kernels
Yanardag, P. and Vishwanathan, S · 2015
Earlier work this paper cites.
Diffusion-convolutional neural networks
Atwood, J. and Towsley, D · 2016
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Convolutional neural networks on graphs with fast localized spectral filtering
Defferrard, M., Bresson, X., and Vandergheynst, P · 2016
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, T. and Kingma, D. P · 2016
Earlier work this paper cites.
Improved techniques for training gans
Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., and Chen, X · 2016
Earlier work this paper cites.
Learning multiagent communication with backpropagation
Sukhbaatar, S., Fergus, R., et al · 2016
Earlier work this paper cites.
Instance normalization: The missing ingredient for fast stylization
Ulyanov, D., Vedaldi, A., and Lempitsky, V · 2016
Earlier work this paper cites.
Neural message passing for quantum chemistry
Gilmer, J., Schoenholz, S. S., Riley, P. F., Vinyals, O., and Dahl, G. E · 2017
Earlier work this paper cites.
Inductive representation learning on large graphs
Hamilton, W., Ying, Z., and Leskovec, J · 2017
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Kipf, T. N. and Welling, M · 2017
Cited alongside, same era.
Geometric deep learning on graphs and manifolds using mixture model cnns
Monti, F., Boscaini, D., Masci, J., Rodola, E., Svoboda, J., and Bronstein, M. M · 2017
Cited alongside, same era.
A simple neural network module for relational reasoning
Santoro, A., Raposo, D., Barrett, D. G., Malinowski, M., Pascanu, R., Battaglia, P., and Lillicrap, T · 2017
Cited alongside, same era.
Moleculenet: A benchmark for molecular machine learning
Wu, Z., Ramsundar, B., Feinberg, E. N., Gomes, J., Geniesse, C., Pappu, A. S., Leswing, K., and Pande, V. S · 2017
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A · 2018
Cited alongside, same era.
Approximation ratios of graph neural networks for combinatorial problems
Sato, R., Yamada, M., and Kashima, H · 2019
Later among the works it cites.
High-dimensional statistics: A non-asymptotic viewpoint , volume 48
Wainwright, M. J · 2019
Later among the works it cites.
How powerful are graph neural networks?
Xu, K., Hu, W., Leskovec, J., and Jegelka, S · 2019
Later among the works it cites.
Towards stabilizing batch statistics in backward propagation of batch normalization
Yan, J., Wan, R., Zhang, X., Zhang, W., Wei, Y., and Sun, J · 2019
Later among the works it cites.
Layer-dependent importance sampling for training deep and large graph convolutional networks
Zou, D., Hu, Z., Wang, Y., Jiang, S., Sun, Y., and Gu, Q · 2019
Later among the works it cites.
Learning graph normalization for graph neural networks, 2020
Chen, Y., Tang, X., Qi, X., Li, C.-G., and Xiao, R · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Norm matters: efficient and accurate normalization schemes in deep networks
Hoffer, E., Banner, R., Golan, I., and Soudry, D · 2018
Cited alongside, same era.
Anonymous walk embeddings
Ivanov, S. and Burnaev, E · 2018
Cited alongside, same era.
Spectral normalization for generative adversarial networks
Miyato, T., Kataoka, T., Koyama, M., and Yoshida, Y · 2018
Cited alongside, same era.
How does batch normalization help optimization?
Santurkar, S., Tsipras, D., Ilyas, A., and Madry, A · 2018
Cited alongside, same era.
The vapnik–chervonenkis dimension of graph and recursive neural networks
Scarselli, F., Tsoi, A. C., and Hagenbuchner, M · 2018
Cited alongside, same era.
Graph attention networks
Velickovic, P., Cucurull, G., Casanova, A., Romero, A., Lio, P., and Bengio, Y · 2018
Cited alongside, same era.
Group normalization
Wu, Y. and He, K · 2018
Cited alongside, same era.
Closest in time.
Benchmarking graph neural networks
Dwivedi, V. P., Joshi, C. K., Laurent, T., Bengio, Y., and Bresson, X · 2020
Closest in time.
Open graph benchmark: Datasets for machine learning on graphs
Hu, W., Fey, M., Zitnik, M., Dong, Y., Ren, H., Liu, B., Catasta, M., and Leskovec, J · 2020
Closest in time.
Deepergcn: All you need to train deeper gcns
Li, G., Xiong, C., Thabet, A., and Ghanem, B · 2020
Closest in time.
How hard is to distinguish graphs with graph neural networks?
Loukas, A · 2020
Closest in time.
Powernorm: Rethinking batch normalization in transformers, 2020
Shen, S., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K · 2020
Closest in time.
A deep learning approach to antibiotic discovery
Stokes, J. M., Yang, K., Swanson, K., Jin, W., Cubillos-Ruiz, A., Donghia, N. M., MacNair, C. R., French, S., Carfrae, L. A., Bloom-Ackerman, Z., et al · 2020
Closest in time.
Acne: Attentive context normalization for robust permutation-equivariant learning
Sun, W., Jiang, W., Trulls, E., Tagliasacchi, A., and Yi, K. M · 2020
Closest in time.
A comprehensive survey on graph neural networks
Wu, Z., Pan, S., Chen, F., Long, G., Zhang, C., and Philip, S. Y · 2020
Closest in time.
On layer normalization in the transformer architecture
Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T.-Y · 2020
Closest in time.
What can neural networks reason about?
Xu, K., Li, J., Zhang, M., Du, S. S., ichi Kawarabayashi, K., and Jegelka, S · 2020
Closest in time.
Revisiting” over-smoothing” in deep gcns
Yang, C., Wang, R., Yao, S., Liu, S., and Abdelzaher, T · 2020
Closest in time.
Deep learning on graphs: A survey
Zhang, Z., Cui, P., and Zhu, W · 2020
Closest in time.
Pairnorm: Tackling oversmoothing in gnns
Zhao, L. and Akoglu, L · 2020
Closest in time.
Gradient descent optimizes over-parameterized deep relu networks
Zou, D., Cao, Y., Zhou, D., and Gu, Q · 2020
Closest in time.
How neural networks extrapolate: From feedforward to graph neural networks
Xu, K., Zhang, M., Li, J., Du, S. S., Kawarabayashi, K.-I., and Jegelka, S · 2021
Closest in time.
Do transformers really perform bad for graph representation?, 2021
Ying, C., Cai, T., Luo, S., Zheng, S., Ke, G., He, D., Shen, Y., and Liu, T.-Y · 2021
Closest in time.