Fetching the paper…
Reading the bibliography…
In this paper, we show that structures similar to self-attention are natural to learn many sequence-to-sequence problems from the perspective of symmetry.
The classical groups: their invariants and representations
Hermann Weyl · 1946
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
Yann LeCun, Bernhard Boser, John S Denker, Donnie Henderson, Richard E Howard, Wayne Hubbard, and Lawrence D Jackel · 1989
Earlier work this paper cites.
Incorporating symmetry into deep dynamics models for improved generalization
Rui Wang, Robin Walters, and Rose Yu · 2002
Earlier work this paper cites.
Lie groups: an approach through invariants and representations
Claudio Procesi · 2006
Earlier work this paper cites.
Why are convolutional nets more sample-efficient than fully-connected nets?
Zhiyuan Li, Yi Zhang, and Sanjeev Arora · 2010
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D Manning · 2015
Earlier work this paper cites.
Permutation-equivariant neural networks applied to dynamics prediction
Nicholas Guttenberg, Nathaniel Virgo, Olaf Witkowski, Hidetoshi Aoki, and Ryota Kanai · 2016
Earlier work this paper cites.
A decomposable attention model for natural language inference
Ankur P Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit · 2016
Earlier work this paper cites.
A structured self-attentive sentence embedding
Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio · 2017
Earlier work this paper cites.
A deep reinforced model for abstractive summarization
Romain Paulus, Caiming Xiong, and Richard Socher · 2017
Earlier work this paper cites.
Equivariance through parameter-sharing
Siamak Ravanbakhsh, Jeff Schneider, and Barnabas Poczos · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Deep sets
Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Russ R Salakhutdinov, and Alexander J Smola · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2018
Cited alongside, same era.
Tensor field networks: Rotation-and translation-equivariant neural networks for 3d point clouds
Nathaniel Thomas, Tess Smidt, Steven Kearnes, Lusann Yang, Li Li, Kai Kohlhoff, and Patrick Riley · 2018
Cited alongside, same era.
Deep potential molecular dynamics: a scalable model with the accuracy of quantum mechanics
Linfeng Zhang, Jiequn Han, Han Wang, Roberto Car, and EJPRL Weinan · 2018
On the sample complexity of learning with geometric stability
Alberto Bietti, Luca Venturi, and Joan Bruna · 2021
Later among the works it cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Later among the works it cites.
Skyformer: Remodel self-attention with gaussian kernel and nyström method
Yifan Chen, Qi Zeng, Heng Ji, and Yun Yang · 2021
Later among the works it cites.
Provably strict generalisation benefit for equivariant models
Bryn Elesedy and Sheheryar Zaidi · 2021
Later among the works it cites.
Learning with invariances in random features and kernel models
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Cormorant: Covariant molecular neural networks
Brandon Anderson, Truong Son Hy, and Risi Kondor · 2019
Cited alongside, same era.
Rotation equivariant and invariant neural networks for microscopy image analysis
Benjamin Chidester, Tianming Zhou, Minh N Do, and Jian Ma · 2019
Cited alongside, same era.
Physical symmetries embedded in neural networks
Marios Mattheakis, Pavlos Protopapas, David Sondak, Marco Di Giovanni, and Efthimios Kaxiras · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Se (3)-transformers: 3d roto-translation equivariant attention networks
Fabian Fuchs, Daniel Worrall, Volker Fischer, and Max Welling · 2020
Cited alongside, same era.
Cycnn: A rotation invariant cnn using polar mapping and cylindrical convolution layers
Jinpyo Kim, Wooekun Jung, Hyungmo Kim, and Jaejin Lee · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al · 2020
Cited alongside, same era.
Later among the works it cites.
A review on the attention mechanism of deep learning
Zhaoyang Niu, Guoqiang Zhong, and Hui Yu · 2021
Later among the works it cites.
A permutation-equivariant neural network architecture for auction design
Jad Rahme, Samy Jelassi, Joan Bruna, and S Matthew Weinberg · 2021
Later among the works it cites.
Kernel self-attention for weakly-supervised image classification using deep multiple instance learning
Dawid Rymarczyk, Adriana Borowa, Jacek Tabor, and Bartosz Zieliński · 2021
Later among the works it cites.
E (n) equivariant graph neural networks
Vıctor Garcia Satorras, Emiel Hoogeboom, and Max Welling · 2021
Later among the works it cites.
Equivariant message passing for the prediction of tensorial properties and molecular spectra
Kristof Schütt, Oliver Unke, and Michael Gastegger · 2021
Later among the works it cites.
Rotation invariant graph neural networks using spin convolutions
Muhammed Shuaibi, Adeesh Kolluru, Abhishek Das, Aditya Grover, Anuroop Sriram, Zachary Ulissi, and C Lawrence Zitnick · 2021
Later among the works it cites.
Scalars are universal: Equivariant machine learning, structured like classical physics
Soledad Villar, David W Hogg, Kate Storey-Fisher, Weichi Yao, and Ben Blum-Smith · 2021
Later among the works it cites.