Fetching the paper…
Reading the bibliography…
The Transformer architecture has become a dominant choice in many domains, such as natural language processing and computer vision.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 1901
Earlier work this paper cites.
Laplacian eigenmaps for dimensionality reduction and data representation
Mikhail Belkin and Partha Niyogi · 2003
Earlier work this paper cites.
The promotion and presentation of the self: celebrity as marker of presentational media
P David Marshall · 2010
Earlier work this paper cites.
To see and be seen: Celebrity practice on twitter
Alice Marwick and Danah Boyd · 2011
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling · 2016
Earlier work this paper cites.
Xavier Bresson and Thomas Laurent · 2017
Earlier work this paper cites.
Neural message passing for quantum chemistry
Justin Gilmer, Samuel S Schoenholz, Patrick F Riley, Oriol Vinyals, and George E Dahl · 2017
Earlier work this paper cites.
Inductive representation learning on large graphs
William L Hamilton, Zhitao Ying, and Jure Leskovec · 2017
Earlier work this paper cites.
Learning graph-level representation for drug discovery
Junying Li, Deng Cai, and Xiaofei He · 2017
Earlier work this paper cites.
Pubchemqc project: a large-scale first-principles electronic structure database for data-driven chemistry
Maho Nakata and Tomomi Shimazaki · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Multi-hop knowledge graph reasoning with reward shaping
Xi Victoria Lin, Richard Socher, and Caiming Xiong · 2018
Earlier work this paper cites.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2018
Earlier work this paper cites.
Graph attention networks
Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio · 2018
Earlier work this paper cites.
Path-augmented graph transformer network
Benson Chen, Regina Barzilay, and Tommi Jaakkola · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Exploiting edge features for graph neural networks
Liyu Gong and Qiang Cheng · 2019
Earlier work this paper cites.
Global relational models of source code
Vincent J Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis, and David Bieber · 2019
Earlier work this paper cites.
Katsuhiko Ishiguro, Shin-ichi Maeda, and Masanori Koyama · 2019
Earlier work this paper cites.
Graph transformer, 2019
Yuan Li, Xiaodan Liang, Zhiting Hu, Yinbo Chen, and Eric P. Xing · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Provably powerful graph networks
Haggai Maron, Heli Ben-Hamu, Hadar Serviansky, and Yaron Lipman · 2019
Cited alongside, same era.
How powerful are graph neural networks?
Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka · 2019
Cited alongside, same era.
Spagan: Shortest path graph attention network
Yiding Yang, Xinchao Wang, Mingli Song, Junsong Yuan, and Dacheng Tao · 2019
Cited alongside, same era.
Position-aware graph neural networks
Jiaxuan You, Rex Ying, and Jure Leskovec · 2019
Cited alongside, same era.
Graph transformer networks
Seongjun Yun, Minbyul Jeong, Raehyun Kim, Jaewoo Kang, and Hyunwoo J Kim · 2019
Cited alongside, same era.
Graph convolutions that can finally model local structure
Rémy Brossard, Oriel Frigo, and David Dehaene · 2020
Cited alongside, same era.
Masked label prediction: Unified message passing model for semi-supervised classification
Yunsheng Shi, Zhengjie Huang, Wenjin Wang, Hui Zhong, Shikun Feng, and Yu Sun · 2020
Later among the works it cites.
Direct multi-hop attention based graph neural network
Guangtao Wang, Rex Ying, Jing Huang, and Jure Leskovec · 2020
Later among the works it cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Later among the works it cites.
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu · 2020
Later among the works it cites.
Breaking the expressive bottlenecks of graph neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Graph transformer for graph-to-sequence learning
Deng Cai and Wai Lam · 2020
Cited alongside, same era.
Principal neighbourhood aggregation for graph nets
Gabriele Corso, Luca Cavalleri, Dominique Beaini, Pietro Liò, and Petar Veličković · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Benchmarking graph neural networks
Vijay Prakash Dwivedi, Chaitanya K Joshi, Thomas Laurent, Yoshua Bengio, and Xavier Bresson · 2020
Cited alongside, same era.
Conformer: Convolution-augmented transformer for speech recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al · 2020
Cited alongside, same era.
Strategies for pre-training graph neural networks
W Hu, B Liu, J Gomes, M Zitnik, P Liang, V Pande, and J Leskovec · 2020
Cited alongside, same era.
Mingqi Yang, Yanming Shen, Heng Qi, and Baocai Yin · 2020
Later among the works it cites.
Graph-bert: Only attention is needed for learning graph representations
Jiawei Zhang, Haopeng Zhang, Congying Xia, and Li Sun · 2020
Later among the works it cites.
Freelb: Enhanced adversarial training for natural language understanding
Chen Zhu, Yu Cheng, Zhe Gan, Siqi Sun, Tom Goldstein, and Jingjing Liu · 2020
Later among the works it cites.
Language-agnostic representation learning of source code from structure and context
Daniel Zügner, Tobias Kirschstein, Michele Catasta, Jure Leskovec, and Stephan Günnemann · 2020
Later among the works it cites.
Accurate learning of graph representations with graph multiset pooling
Jinheon Baek, Minki Kang, and Sung Ju Hwang · 2021
Closest in time.
Directional graph networks
Dominique Beaini, Saro Passaro, Vincent Létourneau, William L Hamilton, Gabriele Corso, and Pietro Liò · 2021
Closest in time.
Graphnorm: A principled approach to accelerating graph neural network training
Tianle Cai, Shengjie Luo, Keyulu Xu, Di He, Tie-yan Liu, and Liwei Wang · 2021
Closest in time.
A generalization of transformer networks to graphs
Vijay Prakash Dwivedi and Xavier Bresson · 2021
Closest in time.
Ogb-lsc: A large-scale challenge for machine learning on graphs
Weihua Hu, Matthias Fey, Hongyu Ren, Maho Nakata, Yuxiao Dong, and Jure Leskovec · 2021
Closest in time.
Rethinking graph transformers with spectral attention
Devin Kreuzer, Dominique Beaini, William Hamilton, Vincent Létourneau, and Prudencio Tossou · 2021
Closest in time.
Parameterized hypercomplex graph neural networks for graph classification
Tuan Le, Marco Bertolini, Frank Noé, and Djork-Arné Clevert · 2021
Closest in time.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Closest in time.
Stable, fast and accurate: Kernelized attention with relative positional encoding
Shengjie Luo, Shanda Li, Tianle Cai, Di He, Dinglan Peng, Shuxin Zheng, Guolin Ke, Liwei Wang, and Tie-Yan Liu · 2021
Closest in time.
Do transformer modifications transfer across implementations and applications?
Sharan Narang, Hyung Won Chung, Yi Tay, William Fedus, Thibault Fevry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam Shazeer, Zhenzhong Lan, et al · 2021
Closest in time.
How could neural networks understand programs?
Dinglan Peng, Shuxin Zheng, Yatao Li, Guolin Ke, Di He, and Tie-Yan Liu · 2021
Closest in time.
Lazyformer: Self attention with lazy update
Chengxuan Ying, Guolin Ke, Di He, and Tie-Yan Liu · 2021
Closest in time.
First place solution of kdd cup 2021 & ogb large-scale challenge graph-level track
Chengxuan Ying, Mingqi Yang, Shuxin Zheng, Guolin Ke, Shengjie Luo, Tianle Cai, Chenglin Wu, Yuxin Wang, Yanming Shen, and Di He · 2021
Closest in time.