Fetching the paper…
Reading the bibliography…
This article derives and validates three principles for initialization and architecture selection in finite width graph neural networks (GNNs) with ReLU activations.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi · 1994
Earlier work this paper cites.
Limit of the smallest eigenvalue of a large dimensional sample covariance matrix
Zhi-Dong Bai and Yong-Qua Yin · 2008
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Exact recovery in the stochastic block model
Emmanuel Abbe, Afonso S. Bandeira, and Georgina Hall · 2016
Earlier work this paper cites.
Sharp nonasymptotic bounds on the norm of random matrices with independent entries
Afonso S Bandeira and Ramon Van Handel · 2016
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling · 2016
Earlier work this paper cites.
Exponential expressivity in deep neural networks through transient chaos
Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli · 2016
Earlier work this paper cites.
On the expressive power of deep neural networks
Maithra Raghu, Ben Poole, Jon M. Kleinberg, Surya Ganguli, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
http://people.idsia.ch/~juergen/fundamentaldeeplearningproblem.html
Sepp hochreiter’s fundamental deep learning problem (1991) · 2017
Earlier work this paper cites.
Samuel S Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Community detection and stochastic block models: Recent developments
Emmanuel Abbe · 2018
Earlier work this paper cites.
Which neural net architectures give rise to exploding and vanishing gradients?
Boris Hanin · 2018
Earlier work this paper cites.
How to start training: The effect of initialization and architecture
Boris Hanin and David Rolnick · 2018
Earlier work this paper cites.
Deeper insights into graph convolutional networks for semi-supervised learning
Qimai Li, Zhichao Han, and Xiao-Ming Wu · 2018
Earlier work this paper cites.
Lechao Xiao, Yasaman Bahri, Jascha Sohl-Dickstein, Samuel S Schoenholz, and Jeffrey Pennington · 2018
Cited alongside, same era.
Fast graph representation learning with pytorch geometric
Matthias Fey and Jan Eric Lenssen · 2019
Cited alongside, same era.
Graph neural networks exponentially lose expressive power for node classification
Kenta Oono and Taiji Suzuki · 2019
Cited alongside, same era.
Fixup initialization: Residual learning without normalization
Hongyi Zhang, Yann N Dauphin, and Tengyu Ma · 2019
Cited alongside, same era.
Measuring and relieving the over-smoothing problem for graph neural networks from the topological view
Deli Chen, Yankai Lin, Wei Li, Peng Li, Jie Zhou, and Xu Sun · 2020
Towards deepening graph neural networks: A gntk-based optimization perspective
Wei Huang, Yayong Li, Weitao Du, Jie Yin, Richard Yi Da Xu, Ling Chen, and Miao Zhang · 2021
Later among the works it cites.
Training graph neural networks with 1000 layers
Guohao Li, Matthias Müller, Bernard Ghanem, and Vladlen Koltun · 2021
Later among the works it cites.
Deepgcns: Making gcns go as deep as cnns
Guohao Li, Matthias Müller, Guocheng Qian, Itzel Carolina Delgadillo Perez, Abdulellah Abualshour, Ali Kassem Thabet, and Bernard Ghanem · 2021
Later among the works it cites.
Skipnode: On alleviating over-smoothing for deep graph convolutional networks
Weigang Lu, Yibing Zhan, Ziyu Guan, Liu Liu, Baosheng Yu, Wei Zhao, Yaming Yang, and Dacheng Tao · 2021
Later among the works it cites.
New insights into graph convolutional networks using neural tangent kernels
Mahalakshmi Sabanayagam, Pascal Esser, and Debarghya Ghoshdastidar · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A note on over-smoothing for graph neural networks
Chen Cai and Yusu Wang · 2020
Cited alongside, same era.
Simple and deep graph convolutional networks
Ming Chen, Zhewei Wei, Zengfeng Huang, Bolin Ding, and Yaliang Li · 2020
Cited alongside, same era.
Products of many large random matrices and gradients in deep neural networks
Boris Hanin and Mihai Nica · 2020
Cited alongside, same era.
Towards deeper graph neural networks
Meng Liu, Hongyang Gao, and Shuiwang Ji · 2020
Cited alongside, same era.
A comprehensive survey on graph neural networks
Zonghan Wu, Shirui Pan, Fengwen Chen, Guodong Long, Chengqi Zhang, and S Yu Philip · 2020
Cited alongside, same era.
Revisiting over-smoothing in deep gcns
Chaoqi Yang, Ruijie Wang, Shuochao Yao, Shengzhong Liu, and Tarek Abdelzaher · 2020
Cited alongside, same era.
Graph neural networks: A review of methods and applications
Jie Zhou, Ganqu Cui, Shengding Hu, Zhengyan Zhang, Cheng Yang, Zhiyuan Liu, Lifeng Wang, Changcheng Li, and Maosong Sun · 2020
Cited alongside, same era.
Later among the works it cites.
Deeper-gxx: deepening arbitrary gnns
Lecheng Zheng, Dongqi Fu, Ross Maciejewski, and Jingrui He · 2021
Later among the works it cites.
Random fully connected neural networks as perturbatively solvable hierarchies
Boris Hanin · 2022
Later among the works it cites.
Old can be gold: Better gradient flow can make vanilla-gcns great again
Ajay Jaiswal, Peihao Wang, Tianlong Chen, Justin Rousseau, Ying Ding, and Zhangyang Wang · 2022
Later among the works it cites.
Not too little, not too much: a theoretical analysis of graph (over) smoothing
Nicolas Keriven · 2022
Later among the works it cites.
The Principles of Deep Learning Theory: An Effective Theory Approach to Understanding Neural Networks
Daniel A Roberts, Sho Yaida, and Boris Hanin · 2022
Later among the works it cites.
A non-asymptotic analysis of oversmoothing in graph neural networks
Xinyi Wu, Zhengdao Chen, William Wang, and Ali Jadbabaie · 2022
Later among the works it cites.
Two sides of the same coin: Heterophily and oversmoothing in graph convolutional neural networks
Yujun Yan, Milad Hashemi, Kevin Swersky, Yaoqing Yang, and Danai Koutra · 2022
Later among the works it cites.
Effective theory of transformers at initialization
Emily Dinan, Sho Yaida, and Susan Zhang · 2023
Closest in time.
Learning local equivariant representations for large-scale atomistic dynamics
Albert Musaelian, Simon Batzner, Anders Johansson, Lixin Sun, Cameron J Owen, Mordechai Kornbluth, and Boris Kozinsky · 2023
Closest in time.