Fetching the paper…
Reading the bibliography…
We propose an extension to the transformer neural network architecture for general-purpose graph learning by adding a dedicated pathway for pairwise structural information, called edge channels.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
Laplacian eigenmaps and spectral techniques for embedding and clustering.. In Nips , Vol. 14. 585–591
Mikhail Belkin and Partha Niyogi. 2001 · 2001
Earlier work this paper cites.
A new model for learning in graph domains. In Proceedings. 2005 IEEE International Joint Conference on Neural Networks, 2005. , Vol. 2. IEEE, 729–734
Marco Gori, Gabriele Monfardini, and Franco Scarselli. 2005 · 2005
Earlier work this paper cites.
The graph neural network model
Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini. 2008 · 2008
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter. 2015 · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling. 2016 · 2016
Earlier work this paper cites.
Xavier Bresson and Thomas Laurent. 2017 · 2017
Earlier work this paper cites.
Neural message passing for quantum chemistry. In International conference on machine learning . PMLR, 1263–1272
Justin Gilmer, Samuel S Schoenholz, Patrick F Riley, Oriol Vinyals, and George E Dahl. 2017 · 2017
Earlier work this paper cites.
Representation learning on graphs: Methods and applications
William L Hamilton, Rex Ying, and Jure Leskovec. 2017 · 2017
Earlier work this paper cites.
Attention is all you need. In Advances in neural information processing systems . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio. 2017 · 2017
Earlier work this paper cites.
How powerful are graph neural networks?
Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka. 2018 · 2018
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
On the relationship between self-attention and convolutional layers
Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi. 2019 · 2019
Earlier work this paper cites.
Neural speech synthesis with transformer network. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 33. 6706–6713
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu. 2019 · 2019
Cited alongside, same era.
Relational pooling for graph representations. In International Conference on Machine Learning . PMLR, 4663–4673
Ryan Murphy, Balasubramaniam Srinivasan, Vinayak Rao, and Bruno Ribeiro. 2019 · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Cited alongside, same era.
Stand-alone self-attention in vision models
Prajit Ramachandran, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jonathon Shlens. 2019 · 2019
Cited alongside, same era.
On the equivalence between positional node embeddings and structural graph representations
Heterogeneous graph transformer. In Proceedings of The Web Conference 2020 . 2704–2710
Ziniu Hu, Yuxiao Dong, Kuansan Wang, and Yizhou Sun. 2020a · 2020
Later among the works it cites.
Transformers are rnns: Fast autoregressive transformers with linear attention. In International Conference on Machine Learning . PMLR, 5156–5165
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. 2020 · 2020
Later among the works it cites.
Flag: Adversarial data augmentation for graph neural networks
Kezhi Kong, Guohao Li, Mucong Ding, Zuxuan Wu, Chen Zhu, Bernard Ghanem, Gavin Taylor, and Tom Goldstein. 2020 · 2020
Later among the works it cites.
Deepergcn: All you need to train deeper gcns
Guohao Li, Chenxin Xiong, Ali Thabet, and Bernard Ghanem. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Balasubramaniam Srinivasan and Bruno Ribeiro. 2019 · 2019
Cited alongside, same era.
Graph transformer networks
Seongjun Yun, Minbyul Jeong, Raehyun Kim, Jaewoo Kang, and Hyunwoo J Kim. 2019 · 2019
Cited alongside, same era.
Graph convolutions that can finally model local structure
Rémy Brossard, Oriel Frigo, and David Dehaene. 2020 · 2020
Cited alongside, same era.
Graph transformer for graph-to-sequence learning. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 34. 7464–7471
Deng Cai and Wai Lam. 2020 · 2020
Cited alongside, same era.
Generative pretraining from pixels. In International Conference on Machine Learning . PMLR, 1691–1703
Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. 2020 · 2020
Cited alongside, same era.
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, and Others. 2020 · 2020
Cited alongside, same era.
Principal neighbourhood aggregation for graph nets
Gabriele Corso, Luca Cavalleri, Dominique Beaini, Pietro Liò, and Petar Veličković. 2020 · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Omri Puny, Heli Ben-Hamu, and Yaron Lipman. 2020 · 2020
Later among the works it cites.
Self-supervised graph transformer on large-scale molecular data
Yu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie, Ying Wei, Wenbing Huang, and Junzhou Huang. 2020 · 2020
Later among the works it cites.
On the Global Self-attention Mechanism for Graph Convolutional Networks. In 2020 25th International Conference on Pattern Recognition (ICPR) . IEEE, 8531–8538
Chen Wang and Chengyuan Deng. 2021 · 2020
Later among the works it cites.
On layer normalization in the transformer architecture. In International Conference on Machine Learning . PMLR, 10524–10533
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu. 2020 · 2020
Later among the works it cites.
Graph-bert: Only attention is needed for learning graph representations
Jiawei Zhang, Haopeng Zhang, Congying Xia, and Li Sun. 2020b · 2020
Later among the works it cites.
Revisiting graph neural networks for link prediction
Muhan Zhang, Pan Li, Yinglong Xia, Kai Wang, and Long Jin. 2020a · 2020
Later among the works it cites.
Directional graph networks. In International Conference on Machine Learning . PMLR, 748–758
Dominique Beani, Saro Passaro, Vincent Létourneau, Will Hamilton, Gabriele Corso, and Pietro Liò. 2021 · 2021
Closest in time.
Ogb-lsc: A large-scale challenge for machine learning on graphs
Weihua Hu, Matthias Fey, Hongyu Ren, Maho Nakata, Yuxiao Dong, and Jure Leskovec. 2021 · 2021
Closest in time.
Parameterized hypercomplex graph neural networks for graph classification. In International Conference on Artificial Neural Networks . Springer, 204–216
Tuan Le, Marco Bertolini, Frank Noé, and Djork-Arné Clevert. 2021 · 2021
Closest in time.
Linear transformers are secretly fast weight memory systems
Imanol Schlag, Kazuki Irie, and Jürgen Schmidhuber. 2021 · 2021
Closest in time.
Do Transformers Really Perform Bad for Graph Representation?
Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng, Guolin Ke, Di He, Yanming Shen, and Tie-Yan Liu. 2021 · 2021
Closest in time.