Fetching the paper…
Reading the bibliography…
Since the introduction of the transformer model by Vaswani et al.
Recurrent neural network based language model
Tomas Mikolov, M. Karafiát, L. Burget, J. Cernocký, and S. Khudanpur · 2010
Earlier work this paper cites.
Context dependent recurrent neural network language model
Tomas Mikolov and G. Zweig · 2012
Earlier work this paper cites.
Recurrent neural network regularization, 2014
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals · 2014
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
A decomposable attention model for natural language inference
Ankur Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit · 2016
Earlier work this paper cites.
Tying word vectors and word classifiers: A loss framework for language modeling
Hakan Inan, Khashayar Khosravi, and Richard Socher · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli · 2018
Earlier work this paper cites.
Scaling neural machine translation
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli · 2018
Earlier work this paper cites.
A simple method for commonsense reasoning, 2018
Trieu H. Trinh and Quoc V. Le · 2018
Earlier work this paper cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Openwebtext corpus
Aaron Gokaslan and Vanya Cohen · 2019
Earlier work this paper cites.
Music transformer: Generating music with long-term structure
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Ian Simon, Curtis Hawthorne, Noam M. Shazeer, Andrew M. Dai, M. Hoffman, M. Dinculescu, and D. Eck · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach, 2019
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
On the relation between position information and sentence length in neural machine translation
Masato Neishi and Naoki Yoshinaga · 2019
Cited alongside, same era.
Analysis of positional encodings for neural machine translation
Jan Rosendahl, Viet Anh Khoa Tran, Weiyue Wang, and Hermann Ney · 2019
Cited alongside, same era.
Longformer: The long-document transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan · 2020
Cited alongside, same era.
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Shape: Shifted absolute position embedding for transformers
Shun Kiyono, Sosuke Kobayashi, Jun Suzuki, and Kentaro Inui · 2021
Closest in time.
Towards mental time travel: a hierarchical memory for reinforcement learning agents
Andrew Kyle Lampinen, Stephanie C. Y. Chan, Andrea Banino, and Felix Hill · 2021
Closest in time.
Base layers: Simplifying training of large, sparse models, 2021
Mike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal, and Luke Zettlemoyer · 2021
Closest in time.
Jurassic-1: Technical details and evaluation
Opher Lieber, Or Sharir, Barak Lenz, and Yoav Shoham · 2021
Closest in time.
CAPE: encoding relative positions with continuous augmented positional embeddings
Tatiana Likhomanenko, Qiantong Xu, Ronan Collobert, Gabriel Synnaeve, and Alex Rogozhnikov · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov · 2020
Cited alongside, same era.
Compositionality decomposed: How do neural networks generalise?
Dieuwke Hupkes, Verna Dankers, Mathijs Mul, and Elia Bruni · 2020
Cited alongside, same era.
Generalization through Memorization: Nearest Neighbor Language Models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis · 2020
Cited alongside, same era.
The eos decision and length extrapolation
Benjamin Newman, John Hewitt, Percy Liang, and Christopher D. Manning · 2020
Cited alongside, same era.
Improving transformer models by reordering their sublayers
Ofir Press, Noah A. Smith, and Omer Levy · 2020
Cited alongside, same era.
Compressive transformers for long-range sequence modelling
Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier, and Timothy P. Lillicrap · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Cited alongside, same era.
Closest in time.
Do transformer modifications transfer across implementations and applications?, 2021
Sharan Narang, Hyung Won Chung, Yi Tay, William Fedus, Thibault Fevry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam Shazeer, Zhenzhong Lan, Yanqi Zhou, Wei Li, Nan Ding, Jake Marcus, Adam Roberts, and Colin Raffel · 2021
Closest in time.
Investigating the limitations of the transformers with simple arithmetic tasks
Rodrigo Nogueira, Zhiying Jiang, and Jimmy J. Li · 2021
Closest in time.
Shortformer: Better language modeling using shorter inputs
Ofir Press, Noah A. Smith, and Mike Lewis · 2021
Closest in time.
Roformer: Enhanced transformer with rotary position embedding, 2021
Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu · 2021
Closest in time.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki · 2021
Closest in time.
The case for translation-invariant self-attention in transformer-based language models, 2021
Ulme Wennberg and Gustav Eje Henter · 2021
Closest in time.
DA-transformer: Distance-aware transformer
Chuhan Wu, Fangzhao Wu, and Yongfeng Huang · 2021
Closest in time.
Using the output embedding to improve language models
Ofir Press and Lior Wolf · 2025
Closest in time.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2074
Closest in time.