Fetching the paper…
Reading the bibliography…
Transformer-based models show their effectiveness across multiple domains and tasks.
A logical calculus of the ideas immanent in nervous activity
Warren S McCulloch and Walter Pitts · 1943
Earlier work this paper cites.
Kleene. representation of events in nerve nets and finite automata
C Stephen · 1956
Earlier work this paper cites.
Backpropagation through time: what it does and how to do it
Paul J Werbos · 1990
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Large text compression benchmark, 2006
Matt Mahoney · 2006
Earlier work this paper cites.
On the properties of neural machine translation: Encoder–decoder approaches
Kyunghyun Cho, Bart van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio · 2014
Earlier work this paper cites.
Neural turing machines, 2014
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Earlier work this paper cites.
Memory networks, 2014
Jason Weston, Sumit Chopra, and Antoine Bordes · 2014
Earlier work this paper cites.
Learning to transduce with unbounded memory, 2015
Edward Grefenstette, Karl Moritz Hermann, Mustafa Suleyman, and Phil Blunsom · 2015
Earlier work this paper cites.
Inferring algorithmic patterns with stack-augmented recurrent nets, 2015
Armand Joulin and Tomas Mikolov · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
End-to-end memory networks, 2015
Sainbayar Sukhbaatar, Arthur Szlam, Jason Weston, and Rob Fergus · 2015
Earlier work this paper cites.
Using fast weights to attend to the recent past
Jimmy Ba, Geoffrey E Hinton, Volodymyr Mnih, Joel Z Leibo, and Catalin Ionescu · 2016
Earlier work this paper cites.
Hybrid computing using a neural network with dynamic external memory
Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka Grabska-Barwińska, Sergio Gómez Colmenarejo, Edward Grefenstette, Tiago Ramalho, John Agapiou, Adrià Puigdomènech Badia, Karl Moritz Hermann, Yori Zwols, Georg Ostrovski, Adam Cain, Helen King, Christopher Summerfield, Phil Blunsom, Koray Kavukcuoglu, and Demis Hassabis · 2016
Earlier work this paper cites.
Dynamic neural turing machine with soft and hard addressing schemes
Caglar Gulcehre, Sarath Chandar, Kyunghyun Cho, and Yoshua Bengio · 2016
Earlier work this paper cites.
Scaling memory-augmented neural networks with sparse reads and writes, 2016
Jack W Rae, Jonathan J Hunt, Tim Harley, Ivo Danihelka, Andrew Senior, Greg Wayne, Alex Graves, and Timothy P Lillicrap · 2016
Earlier work this paper cites.
Memory augmented neural networks with wormhole connections
Caglar Gulcehre, Sarath Chandar, and Yoshua Bengio · 2017
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Cited alongside, same era.
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Speech-transformer: A no-recurrence sequence-to-sequence model for speech recognition
Linhao Dong, Shuang Xu, and Bo Xu · 2018
Cited alongside, same era.
Context-aware neural model for temporal information extraction
Yuanliang Meng and Anna Rumshisky · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al · 2020
Later among the works it cites.
Addressing some limitations of transformers with feedback memory
Angela Fan, Thibaut Lavril, Edouard Grave, Armand Joulin, and Sainbayar Sukhbaatar · 2020
Later among the works it cites.
Gmat: Global memory augmentation for transformers
Ankit Gupta and Jonathan Berant · 2020
Later among the works it cites.
Mart: Memory-augmented recurrent transformer for coherent video paragraph captioning, 2020
Jie Lei, Liwei Wang, Yelong Shen, Dong Yu, Tamara L. Berg, and Mohit Bansal · 2020
Later among the works it cites.
Linformer: Self-attention with linear complexity, 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Cited alongside, same era.
Generating long sequences with sparse transformers, 2019
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Cited alongside, same era.
Transformer-xl: Attentive language models beyond a fixed-length context, 2019
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V. Le, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Star-transformer, 2019
Qipeng Guo, Xipeng Qiu, Pengfei Liu, Yunfan Shao, Xiangyang Xue, and Zheng Zhang · 2019
Cited alongside, same era.
Semeval-2019 task 4: Hyperpartisan news detection
Johannes Kiesel, Maria Mestre, Rishabh Shukla, Emmanuel Vincent, Payam Adineh, David Corney, Benno Stein, and Martin Potthast · 2019
Cited alongside, same era.
Large memory layers with product keys, 2019
Guillaume Lample, Alexandre Sablayrolles, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou · 2019
Cited alongside, same era.
Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2020
Later among the works it cites.
Memformer: The memory-augmented transformer
Qingyang Wu, Zhenzhong Lan, Jing Gu, and Zhou Yu · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al · 2020
Later among the works it cites.
ERNIE-Doc: A retrospective long-document modeling transformer
SiYu Ding, Junyuan Shang, Shuohuan Wang, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Perceiver io: A general architecture for structured inputs & outputs
Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch, Catalin Ionescu, David Ding, Skanda Koppula, Daniel Zoran, Andrew Brock, Evan Shelhamer, et al · 2021
Later among the works it cites.
Staircase attention for recurrent processing of sequences
Da Ju, Stephen Roller, Sainbayar Sukhbaatar, and Jason Weston · 2021
Later among the works it cites.
∞ \infty -former: Infinite memory transformer
Pedro Henrique Martins, Zita Marinho, and André FT Martins · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Later among the works it cites.
Contrastive document representation learning with graph attention networks
Peng Xu, Xinchi Chen, Xiaofei Ma, Zhiheng Huang, and Bing Xiang · 2021
Later among the works it cites.
Ernie-sparse: Learning hierarchical efficient transformer through regularized self-attention
Yang Liu, Jiaxiang Liu, Li Chen, Yuxiang Lu, Shikun Feng, Zhida Feng, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang · 2022
Closest in time.