Fetching the paper…
Reading the bibliography…
While Transformer-based models have shown impressive language modeling performance, the large computation cost is often prohibitive for practical use.
“Scheduled drophead: A regularization method for transformer models,”
Wangchunshu Zhou, Tao Ge, Furu Wei, Ming Zhou, and Ke Xu, · 1980
Earlier work this paper cites.
“Large text compression benchmark,” http://mattmahoney.net/dc/textdata
Matt Mahoney, · 2009
Earlier work this paper cites.
“Pointer sentinel mixture models,”
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher, · 2016
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Learning sparse neural networks through l_0 regularization,”
Christos Louizos, Max Welling, and Diederik P Kingma, · 2018
Earlier work this paper cites.
“Augmenting self-attention with persistent memory,”
Sainbayar Sukhbaatar, Edouard Grave, Guillaume Lample, Herve Jegou, and Armand Joulin, · 2019
Earlier work this paper cites.
“Transformer-xl: Attentive language models beyond a fixed-length context,”
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G Carbonell, Quoc Le, and Ruslan Salakhutdinov, · 2019
Earlier work this paper cites.
“Adaptive attention span in transformers,”
Sainbayar Sukhbaatar, Édouard Grave, Piotr Bojanowski, and Armand Joulin, · 2019
Cited alongside, same era.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Cited alongside, same era.
“What does bert look at? an analysis of bert’s attention,”
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning, · 2019
Cited alongside, same era.
“Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned,”
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov, · 2019
Cited alongside, same era.
“Are sixteen heads really better than one?,”
Paul Michel, Omer Levy, and Graham Neubig, · 2019
Cited alongside, same era.
“Batch-shaping for learning conditional channel gated networks,”
“Parameter-efficient transfer learning with diff pruning,”
Demi Guo, Alexander M Rush, and Yoon Kim, · 2020
Later among the works it cites.
“Movement pruning: Adaptive sparsity by fine-tuning,”
Victor Sanh, Thomas Wolf, and Alexander Rush, · 2020
Later among the works it cites.
“Self-attention attribution: Interpreting information interactions inside transformer,”
Yaru Hao, Li Dong, Furu Wei, and Ke Xu, · 2021
Closest in time.
“Spatten: Efficient sparse attention architecture with cascade token and head pruning,”
Hanrui Wang, Zhekai Zhang, and Song Han, · 2021
Closest in time.
“Know what you don’t need: Single-shot meta-pruning for attention heads,”
Zhengyan Zhang, Fanchao Qi, Zhiyuan Liu, Qun Liu, and Maosong Sun, · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Babak Ehteshami Bejnordi, Tijmen Blankevoort, and Max Welling, · 2019
Cited alongside, same era.
“Language models are few-shot learners,”
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al., · 2020
Cited alongside, same era.
Shucong Zhang, Erfan Loweimi, Peter Bell, and Steve Renals, · 2021
Closest in time.