Fetching the paper…
Reading the bibliography…
Graph neural networks based on iterative one-hop message passing have been shown to struggle in harnessing the information from distant nodes effectively.
Algorithm 97: shortest path
Floyd, R. W · 1962
Earlier work this paper cites.
A theorem on boolean matrices
Warshall, S · 1962
Earlier work this paper cites.
Prefix sums and their applications, 1990
Blelloch, G. E · 1990
Earlier work this paper cites.
Social influence analysis in large-scale networks
Tang, J., Sun, J., Wang, C., and Yang, Z · 2009
Earlier work this paper cites.
Matrix analysis
Bhatia, R · 2013
Earlier work this paper cites.
Bresson, X. and Laurent, T · 2017
Earlier work this paper cites.
Language modeling with gated convolutional networks
Dauphin, Y. N., Fan, A., Auli, M., and Grangier, D · 2017
Earlier work this paper cites.
Neural message passing for quantum chemistry
Gilmer, J., Schoenholz, S. S., Riley, P. F., Vinyals, O., and Dahl, G. E · 2017
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Kipf, T. N. and Welling, M · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Deep sets
Zaheer, M., Kottur, S., Ravanbakhsh, S., Poczos, B., Salakhutdinov, R. R., and Smola, A. J · 2017
Earlier work this paper cites.
Graph attention networks
Veličković, P., Cucurull, G., Casanova, A., Romero, A., Lio, P., and Bengio, Y · 2018
Earlier work this paper cites.
Graph convolutional neural networks for web-scale recommender systems
Ying, R., He, R., Chen, K., Eksombatchai, P., Hamilton, W. L., and Leskovec, J · 2018
Earlier work this paper cites.
Mixhop: Higher-order graph convolutional architectures via sparsified neighborhood mixing
Abu-El-Haija, S., Perozzi, B., Kapoor, A., Alipourfard, N., Lerman, K., Harutyunyan, H., Ver Steeg, G., and Galstyan, A · 2019
Earlier work this paper cites.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
How powerful are graph neural networks?
Xu, K., Hu, W., Leskovec, J., and Jegelka, S · 2019
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Earlier work this paper cites.
Principal neighbourhood aggregation for graph nets
Corso, G., Cavalleri, L., Beaini, D., Liò, P., and Veličković, P · 2020
Earlier work this paper cites.
Hippo: Recurrent memory with optimal polynomial projections
Gu, A., Dao, T., Ermon, S., Rudra, A., and Ré, C · 2020
Cited alongside, same era.
Reformer: The efficient transformer
Kitaev, N., Kaiser, Ł., and Levskaya, A · 2020
Cited alongside, same era.
Distance encoding: Design provably more powerful neural networks for graph representation learning
Li, P., Wang, Y., Wang, H., and Leskovec, J · 2020
Cited alongside, same era.
k-hop graph neural networks
Nikolentzos, G., Dasoulas, G., and Vazirgiannis, M · 2020
Cited alongside, same era.
Linformer: Self-attention with linear complexity
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2020
Cited alongside, same era.
Lite transformer with long-short range attention
Wu, Z., Liu, Z., Lin, J., Lin, Y., and Han, S · 2020
Pure transformers are powerful graph learners
Kim, J., Nguyen, D., Min, S., Cho, S., Lee, M., Lee, H., and Hong, S · 2022
Later among the works it cites.
Approximation and optimization theory for linear continuous-time recurrent neural networks
Li, Z., Han, J., E, W., and Li, Q · 2022
Later among the works it cites.
S4nd: Modeling images and videos as multidimensional signals using state spaces
Nguyen, E., Goel, K., Gu, A., Downs, G. W., Shah, P., Dao, T., Baccus, S. A., and Ré, C · 2022
Later among the works it cites.
Recipe for a general, powerful, scalable graph transformer
Rampášek, L., Galkin, M., Dwivedi, V. P., Luu, A. T., Wolf, G., and Beaini, D · 2022
Later among the works it cites.
Understanding over-squashing and bottlenecks on graphs via curvature
Topping, J., Di Giovanni, F., Chamberlain, B. P., Dong, X., and Bronstein, M. M · 2022
Later among the works it cites.
On over-squashing in message passing neural networks: The impact of width, depth, and topology
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the bottleneck of graph neural networks and its practical implications
Alon, U. and Yahav, E · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2021
Cited alongside, same era.
Rethinking graph transformers with spectral attention
Kreuzer, D., Beaini, D., Hamilton, W., Létourneau, V., and Tossou, P · 2021
Cited alongside, same era.
Long range arena: A benchmark for efficient transformers
Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D · 2021
Cited alongside, same era.
Mlp-mixer: An all-mlp architecture for vision
Tolstikhin, I. O., Houlsby, N., Kolesnikov, A., Beyer, L., Zhai, X., Unterthiner, T., Yung, J., Steiner, A., Keysers, D., Uszkoreit, J., et al · 2021
Cited alongside, same era.
Representing long-range context for graph neural networks with global attention
Wu, Z., Jain, P., Wright, M., Mirhoseini, A., Gonzalez, J. E., and Stoica, I · 2021
Cited alongside, same era.
Di Giovanni, F., Giusti, L., Barbero, F., Luise, G., Lio, P., and Bronstein, M. M · 2023
Closest in time.
Benchmarking graph neural networks
Dwivedi, V. P., Joshi, C. K., Luu, A. T., Laurent, T., Bengio, Y., and Bresson, X · 2023
Closest in time.
Hungry hungry hippos: Towards language modeling with state space models
Fu, D. Y., Dao, T., Saab, K. K., Thomas, A. W., Rudra, A., and Re, C · 2023
Closest in time.
How to train your hippo: State space models with generalized orthogonal basis projections
Gu, A., Johnson, I., Timalsina, A., Rudra, A., and Ré, C · 2023
Closest in time.
Drew: Dynamically rewired message passing with delay
Gutteridge, B., Dong, X., Bronstein, M. M., and Di Giovanni, F · 2023
Closest in time.
Liquid structural state-space models
Hasani, R., Lechner, M., Wang, T.-H., Chahine, M., Amini, A., and Rus, D · 2023
Closest in time.
A generalization of vit/mlp-mixer to graphs
He, X., Hooi, B., Laurent, T., Perold, A., LeCun, Y., and Bresson, X · 2023
Closest in time.
Graph inductive biases in transformers without message passing
Ma, L., Lin, C., Lim, D., Romero-Soriano, A., Dokania, P. K., Coates, M., Torr, P., and Lim, S.-N · 2023
Closest in time.
Path neural networks: Expressive and accurate graph neural networks
Michel, G., Nikolentzos, G., Lutzeyer, J. F., and Vazirgiannis, M · 2023
Closest in time.
Rwkv: Reinventing rnns for the transformer era
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Cao, H., Cheng, X., Chung, M., Grella, M., GV, K. K., et al · 2023
Closest in time.
Simplified state space layers for sequence modeling
Smith, J. T., Warrington, A., and Linderman, S. W · 2023
Closest in time.
Where did the gap go? reassessing the long-range graph benchmark
Tönshoff, J., Ritzert, M., Rosenbluth, E., and Grohe, M · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Closest in time.
Rethinking the expressive power of gnns via graph biconnectivity
Zhang, B., Luo, S., Wang, L., and He, D · 2023
Closest in time.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2024
Closest in time.