OGB-LSC: A large-scale challenge for machine learning on graphs
W. Hu, M. Fey, H. Ren, M. Nakata, Y. Dong, and J. Leskovec · 2021
Later among the works it cites.
Edge-augmented graph transformers: Global self-attention is enough for graphs
M. S. Hussain, M. J. Zaki, and D. Subramanian · 2021
Later among the works it cites.
Perceiver IO: A general architecture for structured inputs & outputs
A. Jaegle, S. Borgeaud, J. Alayrac, C. Doersch, C. Ionescu, D. Ding, S. Koppula, D. Zoran, A. Brock, E. Shelhamer, O. J. Hénaff, M. M. Botvinick, A. Zisserman, O. Vinyals, and J. Carreira · 2021
Later among the works it cites.
Perceiver: General perception with iterative attention
A. Jaegle, F. Gimeno, A. Brock, O. Vinyals, A. Zisserman, and J. Carreira · 2021
Later among the works it cites.
Transformers generalize deepsets and can be extended to graphs and hypergraphs
J. Kim, S. Oh, and S. Hong · 2021
Later among the works it cites.
Rethinking graph transformers with spectral attention
D. Kreuzer, D. Beaini, W. L. Hamilton, V. Létourneau, and P. Tossou · 2021
Later among the works it cites.
On the expressive power of self-attention matrices
V. Likhosherstov, K. Choromanski, and A. Weller · 2021
Later among the works it cites.
Mesh graphormer
K. Lin, L. Wang, and Z. Liu · 2021
Later among the works it cites.
Luna: Linear unified nested attention
X. Ma, X. Kong, S. Wang, C. Zhou, J. May, H. Ma, and L. Zettlemoyer · 2021
Later among the works it cites.
Random feature attention
H. Peng, N. Pappas, D. Yogatama, R. Schwartz, N. A. Smith, and L. Kong · 2021
Later among the works it cites.
Multi-scale attributed node embedding
B. Rozemberczki, C. Allen, and R. Sarkar · 2021
Later among the works it cites.
Efficient attention: Attention with linear complexities
Z. Shen, M. Zhang, H. Zhao, S. Yi, and H. Li · 2021
Later among the works it cites.
How to train your vit? data, augmentation, and regularization in vision transformers
A. Steiner, A. Kolesnikov, X. Zhai, R. Wightman, J. Uszkoreit, and L. Beyer · 2021
Later among the works it cites.
Which transformer architecture fits my data? A vocabulary bottleneck in self-attention
N. Wies, Y. Levine, D. Jannai, and A. Shashua · 2021
Later among the works it cites.
Nyströmformer: A nyström-based algorithm for approximating self-attention
Y. Xiong, Z. Zeng, R. Chakraborty, M. Tan, G. Fung, Y. Li, and V. Singh · 2021
Later among the works it cites.
Do transformers really perform bad for graph representation?
C. Ying, T. Cai, S. Luo, S. Zheng, G. Ke, D. He, Y. Shen, and T. Liu · 2021
Later among the works it cites.
Incorporating convolution designs into visual transformers
K. Yuan, S. Guo, Z. Liu, A. Zhou, F. Yu, and W. Wu · 2021
Later among the works it cites.
How attentive are graph attention networks?
S. Brody, U. Alon, and E. Yahav · 2022
Closest in time.
Sign and basis invariant networks for spectral graph representation learning
D. Lim, J. Robinson, L. Zhao, T. E. Smidt, S. Sra, H. Maron, and S. Jegelka · 2022
Closest in time.
Transformer for graphs: An overview from architecture perspective
E. Min, R. Chen, Y. Bian, T. Xu, K. Zhao, W. Huang, P. Zhao, J. Huang, S. Ananiadou, and Y. Rong · 2022
Closest in time.
Universal graph transformer self-attention networks
D. Q. Nguyen, T. D. Nguyen, and D. Phung · 2022
Closest in time.
Permutation equivariant layers for higher order interactions
H. Pan and R. Kondor · 2022
Closest in time.
Grpe: Relative positional encoding for graph transformer
W. Park, W. Chang, D. Lee, J. Kim, and S. won Hwang · 2022
Closest in time.
A generalist agent
S. Reed, K. Żolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Giménez, Y. Sulsky, J. Kay, J. T. Springenberg, T. Eccles, J. Bruce, A. Razavi, A. Edwards, N. Heess, Y. Chen, R. Hadsell, O. Vinyals, M. Bordbar, and N. de Freitas · 2022
Closest in time.
Benchmarking graphormer on large-scale molecular modeling datasets
Y. Shi, S. Zheng, G. Ke, Y. Shen, J. You, J. He, S. Luo, C. Liu, D. He, and T. Liu · 2022
Closest in time.
Deepnet: Scaling transformers to 1, 000 layers
H. Wang, S. Ma, L. Dong, S. Huang, D. Zhang, and F. Wei · 2022
Closest in time.
1-wl expressiveness is (almost) all you need
M. Zopf · 2022
Closest in time.