Fetching the paper…
Reading the bibliography…
Transformers flexibly operate over sets of real-valued vectors representing task-specific entities and their attributes, where each vector might encode one word-piece token and its position in a sequence, or some piece of information that carries no position at all.
Structure-activity relationship of mutagenic aromatic and heteroaromatic nitro compounds. correlation with molecular orbital energies and hydrophobicity
Asim Kumar Debnath, Rosa L. Lopez de Compadre, Gargi Debnath, Alan J. Shusterman, and Corwin Hansch · 1991
Earlier work this paper cites.
Automating the construction of internet portals with machine learning
Andrew Kachites McCallum, Kamal Nigam, Jason Rennie, and Kristie Seymore · 2000
Earlier work this paper cites.
Graph-bert: Only attention is needed for learning graph representations, 2020a
Jiawei Zhang, Haopeng Zhang, Congying Xia, and Li Sun · 2001
Earlier work this paper cites.
Signed networks in social media
Jure Leskovec, Daniel Huttenlocher, and Jon Kleinberg · 2010
Earlier work this paper cites.
Translating embeddings for modeling multi-relational data
Antoine Bordes, Nicolas Usunier, Alberto Garcia-Durán, Jason Weston, and Oksana Yakhnenko · 2013
Earlier work this paper cites.
Wojciech Zaremba and Ilya Sutskever · 2014
Earlier work this paper cites.
Neural gpus learn algorithms, 2015
Łukasz Kaiser and Ilya Sutskever · 2015
Earlier work this paper cites.
Neural message passing for quantum chemistry
Justin Gilmer, Samuel S. Schoenholz, Patrick F. Riley, Oriol Vinyals, and George E. Dahl · 2017
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Thomas N. Kipf and Max Welling · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Deep sets
Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Russ R Salakhutdinov, and Alexander J Smola · 2017
Earlier work this paper cites.
Molecular mechanics-driven graph neural network with multiplex graph for molecular structures
Shuo Zhang, Yang Liu, and Lei Xie · 2017
Earlier work this paper cites.
Relational inductive biases, deep learning, and graph networks, 2018
Peter W. Battaglia, Jessica B. Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, Caglar Gulcehre, Francis Song, Andrew Ballard, Justin Gilmer, George Dahl, Ashish Vaswani, Kelsey Allen, Charles Nash, Victoria Langston, Chris Dyer, Nicolas Heess, Daan Wierstra, Pushmeet Kohli, Matt Botvinick, Oriol Vinyals, Yujia Li, and Razvan Pascanu · 2018
Earlier work this paper cites.
Neural arithmetic logic units
Andrew Trask, Felix Hill, Scott E Reed, Jack Rae, Chris Dyer, and Phil Blunsom · 2018
Earlier work this paper cites.
Graph attention networks
Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio · 2018
Earlier work this paper cites.
Moleculenet: a benchmark for molecular machine learning
Zhenqin Wu, Bharath Ramsundar, Evan N. Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S. Pappu, Karl Leswing, and Vijay Pande · 2018
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G. Carbonell, Quoc Viet Le, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Exploiting edge features for graph neural networks
Liyu Gong and Qiang Cheng · 2019
Earlier work this paper cites.
An investigation of model-free planning
Arthur Guez, Mehdi Mirza, Karol Gregor, Rishabh Kabra, Sebastien Racaniere, Theophane Weber, David Raposo, Adam Santoro, Laurent Orseau, Tom Eccles, Greg Wayne, David Silver, and Timothy Lillicrap · 2019
Earlier work this paper cites.
Star-transformer
Qipeng Guo, Xipeng Qiu, Pengfei Liu, Yunfan Shao, Xiangyang Xue, and Zheng Zhang · 2019
Cited alongside, same era.
Censnet: Convolution with edge-node switching in graph neural networks
Xiaodong Jiang, Pengsheng Ji, and Sheng Li · 2019
Cited alongside, same era.
Attention, learn to solve routing problems!
Wouter Kool, Herke van Hoof, and Max Welling · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Global relational models of source code
Vincent J. Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis, and David Bieber · 2020
Cited alongside, same era.
Offline reinforcement learning as one big sequence modeling problem
Michael Janner, Qiyang Li, and Sergey Levine · 2021
Later among the works it cites.
Rethinking graph transformers with spectral attention
Devin Kreuzer, Dominique Beaini, William L. Hamilton, Vincent Létourneau, and Prudencio Tossou · 2021
Later among the works it cites.
Are we really making much progress? revisiting, benchmarking and refining heterogeneous graph neural networks
Qingsong Lv, Ming Ding, Qiang Liu, Yuxiang Chen, Wenzheng Feng, Siming He, Chang Zhou, Jianguo Jiang, Yuxiao Dong, and Jie Tang · 2021
Later among the works it cites.
Representing long-range context for graph neural networks with global attention
Zhanghao Wu, Paras Jain, Matthew Wright, Azalia Mirhoseini, Joseph E Gonzalez, and Ion Stoica · 2021
Later among the works it cites.
Do transformers really perform badly for graph representation?
Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng, Guolin Ke, Di He, Yanming Shen, and Tie-Yan Liu · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ziniu Hu, Yuxiao Dong, Kuansan Wang, and Yizhou Sun · 2020
Cited alongside, same era.
Working memory graphs
Ricky Loynd, Roland Fernandez, Asli Celikyilmaz, Adith Swaminathan, and Matthew Hausknecht · 2020
Cited alongside, same era.
Stabilizing transformers for reinforcement learning
Emilio Parisotto, Francis Song, Jack Rae, Razvan Pascanu, Caglar Gulcehre, Siddhant Jayakumar, Max Jaderberg, Raphaël Lopez Kaufman, Aidan Clark, Seb Noury, Matthew Botvinick, Nicolas Heess, and Raia Hadsell · 2020
Cited alongside, same era.
Towards scale-invariant graph-related problem solving by iterative homogeneous graph neural networks
Hao Tang, Zhiao Huang, Jiayuan Gu, Bao-Liang Lu, and Hao Su · 2020
Cited alongside, same era.
Pointer graph networks
Petar Veličković, Lars Buesing, Matthew Overlan, Razvan Pascanu, Oriol Vinyals, and Charles Blundell · 2020
Cited alongside, same era.
Neural execution of graph algorithms
Petar Veličković, Rex Ying, Matilde Padovano, Raia Hadsell, and Charles Blundell · 2020
Cited alongside, same era.
What can neural networks reason about?
Keyulu Xu, Jingling Li, Mozhi Zhang, Simon S. Du, Ken ichi Kawarabayashi, and Stefanie Jegelka · 2020
Cited alongside, same era.
How attentive are graph attention networks?
Shaked Brody, Uri Alon, and Eran Yahav · 2022
Closest in time.
Structure-aware transformer for graph representation learning
Dexiong Chen, Leslie O’Bray, and Karsten Borgwardt · 2022
Closest in time.
Graph neural networks are dynamic programmers, 2022
Andrew Dudzik and Petar Veličković · 2022
Closest in time.
Heterogeneity-aware twitter bot detection with relational graph transformers
Shangbin Feng, Zhaoxuan Tan, Rui Li, and Minnan Luo · 2022
Closest in time.
Global self-attention as a replacement for graph convolution
Md Shamim Hussain, Mohammed J. Zaki, and Dharmashankar Subramanian · 2022
Closest in time.
A generalist neural algorithmic learner
Borja Ibarz, Vitaly Kurin, George Papamakarios, Kyriacos Nikiforou, Mehdi Bennani, Róbert Csordás, Andrew Joseph Dudzik, Matko Bošnjak, Alex Vitvitskyi, Yulia Rubanova, Andreea Deac, Beatrice Bevilacqua, Yaroslav Ganin, Charles Blundell, and Petar Veličković · 2022
Closest in time.
Pure transformers are powerful graph learners
Jinwoo Kim, Dat Tien Nguyen, Seonwoo Min, Sungjun Cho, Moontae Lee, Honglak Lee, and Seunghoon Hong · 2022
Closest in time.
Mask and reason: Pre-training knowledge graph transformers for complex logical queries
Xiao Liu, Shiyu Zhao, Kai Su, Yukuo Cen, Jiezhong Qiu, Mengdi Zhang, Wei Wu, Yuxiao Dong, and Jie Tang · 2022
Closest in time.
Recipe for a general, powerful, scalable graph transformer
Ladislav Rampasek, Mikhail Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu, Guy Wolf, and Dominique Beaini · 2022
Closest in time.
The CLRS algorithmic reasoning benchmark
Petar Veličković, Adrià Puigdomènech Badia, David Budden, Razvan Pascanu, Andrea Banino, Misha Dashevskiy, Raia Hadsell, and Charles Blundell · 2022
Closest in time.
Nodeformer: A scalable graph structure learning transformer for node classification
Qitian Wu, Wentao Zhao, Zenan Li, David Wipf, and Junchi Yan · 2022
Closest in time.
Reformer: The relational transformer for image captioning
Xuewen Yang, Yingru Liu, and Xin Wang · 2022
Closest in time.
Edgeformers: Graph-empowered transformers for representation learning on textual-edge networks
Bowen Jin, Yu Zhang, Yu Meng, and Jiawei Han · 2023
Closest in time.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2074
Closest in time.