Fetching the paper…
Reading the bibliography…
The graph convolutional network (GCN) is a go-to solution for machine learning on graphs, but its training is notoriously difficult to scale both in terms of graph size and the number of model parameters.
Applications of Graph Theory in Chemistry
Balaban, A. T · 1985
Earlier work this paper cites.
Random graph models of social networks
Newman, M. E., Watts, D. J., and Strogatz, S. H · 2002
Earlier work this paper cites.
A graph-based toy model of chemistry
Benkö, G., Flamm, C., and Stadler, P. F · 2003
Earlier work this paper cites.
A new model for learning in graph domains
Gori, M., Monfardini, G., and Scarselli, F · 2005
Earlier work this paper cites.
Collective classification in network data
Sen, P., Namata, G., Bilgic, M., Getoor, L., Galligher, B., and Eliassi-Rad, T · 2008
Earlier work this paper cites.
Parallelized stochastic gradient descent
Zinkevich, M., Weimer, M., Li, L., and Smola, A. J · 2010
Earlier work this paper cites.
Distributed delayed stochastic optimization
Agarwal, A. and Duchi, J. C · 2011
Earlier work this paper cites.
Exponential random graph models for social networks: Theory, methods, and applications
Lusher, D., Koskinen, J., and Robins, G · 2013
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Geodesic convolutional neural networks on riemannian manifolds
Masci, J., Boscaini, D., Bronstein, M., and Vandergheynst, P · 2015
Earlier work this paper cites.
Deep learning with elastic averaging sgd
Zhang, S., Choromanska, A. E., and LeCun, Y · 2015
Earlier work this paper cites.
Convolutional neural networks on graphs with fast localized spectral filtering
Defferrard, M., Bresson, X., and Vandergheynst, P · 2016
Earlier work this paper cites.
Semi-Supervised Classification with Graph Convolutional Networks
Kipf, T. N. and Welling, M · 2016
Earlier work this paper cites.
Geometric Deep Learning: Going beyond Euclidean data
Bronstein, M. M., Bruna, J., LeCun, Y., Szlam, A., and Vandergheynst, P · 2017
Earlier work this paper cites.
Integrated Model, Batch and Domain Parallelism in Training Neural Networks
Gholami, A., Azad, A., Jin, P., Keutzer, K., and Buluc, A · 2017
Earlier work this paper cites.
Inductive representation learning on large graphs
Hamilton, W., Ying, Z., and Leskovec, J · 2017
Earlier work this paper cites.
Can Decentralized Algorithms Outperform Centralized Algorithms? A Case Study for Decentralized Parallel Stochastic Gradient Descent
Lian, X., Zhang, C., Zhang, H., Hsieh, C.-J., Zhang, W., and Liu, J · 2017
Cited alongside, same era.
Veličković, P., Cucurull, G., Casanova, A., Romero, A., Lio, P., and Bengio, Y · 2017
Cited alongside, same era.
Large-Scale Learnable Graph Convolutional Networks
Gao, H., Wang, Z., and Ji, S · 2018
Cited alongside, same era.
Layer-Parallel Training of Deep Residual Neural Networks
Günther, S., Ruthotto, L., Schroder, J. B., Cyr, E. C., and Gauger, N. R · 2018
Cited alongside, same era.
Adaptive Sampling Towards Fast Graph Representation Learning
Local SGD converges fast and communicates little
Stich, S. U · 2019
Later among the works it cites.
Automatic Model Parallelism for Deep Neural Networks with Compiler and Hardware Support
Tavarageri, S., Sridharan, S., and Kaul, B · 2019
Later among the works it cites.
Layered sgd: A decentralized and synchronous sgd algorithm for scalable deep neural network training
Yu, K., Flynn, T., Yoo, S., and D’Imperio, N · 2019
Later among the works it cites.
Distributed Learning of Deep Neural Networks using Independent Subnet Training
Yuan, B., Kyrillidis, A., and Jermaine, C. M · 2019
Later among the works it cites.
GraphSAINT: Graph Sampling Based Inductive Learning Method
Zeng, H., Zhou, H., Srivastava, A., Kannan, R., and Prasanna, V · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Huang, W., Zhang, T., Rong, Y., and Huang, J · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Cited alongside, same era.
Deeper insights into graph convolutional networks for semi-supervised learning
Li, Q., Han, Z., and Wu, X.-M · 2018
Cited alongside, same era.
Don’t Use Large Mini-Batches, Use Local SGD
Lin, T., Stich, S. U., Kshitij Patel, K., and Jaggi, M · 2018
Cited alongside, same era.
A quick survey on large scale distributed deep learning systems
Zhang, Z., Yin, L., Peng, Y., and Li, D · 2018
Cited alongside, same era.
Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
Ben-Nun, T. and Hoefler, T · 2019
Cited alongside, same era.
Cluster-gcn: An efficient algorithm for training deep and large graph convolutional networks
Chiang, W.-L., Liu, X., Si, S., Li, Y., Bengio, S., and Hsieh, C.-J · 2019
Cited alongside, same era.
Unsupervised Cross-lingual Representation Learning at Scale
Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., Grave, E., Ott, M., Zettlemoyer, L., and Stoyanov, V · 2019
Cited alongside, same era.
Later among the works it cites.
Layer-Dependent Importance Sampling for Training Deep and Large Graph Convolutional Networks
Zou, D., Hu, Z., Wang, Y., Jiang, S., Sun, Y., and Gu, Q · 2019
Later among the works it cites.
Language models are few-shot learners
Brown, T. B. et al · 2020
Later among the works it cites.
Open Graph Benchmark: Datasets for Machine Learning on Graphs
Hu, W., Fey, M., Zitnik, M., Dong, Y., Ren, H., Liu, B., Catasta, M., and Leskovec, J · 2020
Later among the works it cites.
Kirby, A. C., Samsi, S., Jones, M., Reuther, A., Kepner, J., and Gadepally, V · 2020
Later among the works it cites.
Convolutional Neural Network Training with Distributed K-FAC
Pauloski, J. G., Zhang, Z., Huang, L., Xu, W., and Foster, I. T · 2020
Later among the works it cites.
The cost of training nlp models: A concise overview
Sharir, O., Peleg, B., and Shoham, Y · 2020
Later among the works it cites.
A Quantitative Survey of Communication Optimizations in Distributed Deep Learning
Shi, S., Tang, Z., Chu, X., Liu, C., Wang, W., and Li, B · 2020
Later among the works it cites.
Quadratic suffices for over-parametrization via matrix chernoff bound, 2020
Song, Z. and Yang, X · 2020
Later among the works it cites.
L2-gcn: Layer-wise and learned efficient training of graph convolutional networks
You, Y., Chen, T., Wang, Z., and Shen, Y · 2020
Later among the works it cites.
LAMP: Large Deep Nets with Automated Model Parallelism for Image Segmentation
Zhu, W., Zhao, C., Li, W., Roth, H., Xu, Z., and Xu, D · 2020
Later among the works it cites.
On the convergence of shallow neural network training with randomly masked neurons, 2021
Liao, F. and Kyrillidis, A · 2021
Closest in time.