Fetching the paper…
Reading the bibliography…
Multilayer transformer networks consist of interleaved self-attention and feedforward sublayers.
Understanding and improving transformer from a multi-particle dynamic system point of view
Yiping Lu, Zhuohan Li, Di He, Zhiqing Sun, Bin Dong, Tao Qin, Liwei Wang, and Tie-Yan Liu. 2019 · 1906
Earlier work this paper cites.
Improving deep transformer with depth-scaled initialization and merged attention
Biao Zhang, Ivan Titov, and Rico Sennrich. 2019 · 1908
Earlier work this paper cites.
Adaptively sparse transformers
Gonçalo M. Correia, Vlad Niculae, and André F. T. Martins. 2019 · 1909
Earlier work this paper cites.
Transformers without tears: Improving the normalization of self-attention
Toan Q. Nguyen and Julian Salazar. 2019 · 1910
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2019 · 1911
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V. Le. 2016 · 1911
Earlier work this paper cites.
The hungarian method for the assignment problem
Harold W. Kuhn. 1955 · 1955
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
An empirical exploration of recurrent network architectures
Rafal Jozefowicz, Wojciech Zaremba, and Ilya Sutskever. 2015 · 2015
Earlier work this paper cites.
Yukun Zhu, Ryan Kiros, Richard Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016 · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Tying word vectors and word classifiers: A loss framework for language modeling
Hakan Inan, Khashayar Khosravi, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Using the output embedding to improve language models
Ofir Press and Lior Wolf. 2017 · 2017
Cited alongside, same era.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019 · 2019
Closest in time.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Closest in time.
Star-transformer
Qipeng Guo, Xipeng Qiu, Pengfei Liu, Yunfan Shao, Xiangyang Xue, and Zheng Zhang. 2019 · 2019
Closest in time.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Closest in time.
Smaller, faster, cheaper, lighter: Introducing DistilBERT, a distilled version of BERT
Victor Sanh. 2019 · 2019
Closest in time.
The evolved transformer
David So, Quoc Le, and Chen Liang. 2019 · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Scaling neural machine translation
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli. 2018 · 2018
Cited alongside, same era.
Improving language understanding with unsupervised learning
Alec Radford, Karthik Narasimhan, Time Salimans, and Ilya Sutskever. 2018 · 2018
Cited alongside, same era.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli. 2019 · 2019
Cited alongside, same era.
Adaptive attention span in transformers
Sainbayar Sukhbaatar, Edouard Grave, Piotr Bojanowski, and Armand Joulin. 2019 · 2019
Closest in time.
EfficientNet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc Le. 2019 · 2019
Closest in time.
Compressive transformers for long-range sequence modelling
Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier, and Timothy P. Lillicrap. 2020 · 2020
Closest in time.