Fetching the paper…
Reading the bibliography…
How much information do NLP tasks really need from a transformer's attention mechanism at application-time (inference)? From recent work, we know that there is sparsity in transformers and that the floating-points within its computation can be discretized to fewer values with minimal loss to task accuracies.
Generating Long Sequences with Sparse Transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 1904
Earlier work this paper cites.
Efficient 8-Bit Quantization of Transformer Neural Machine Language Translation Model
Aishwarya Bhandare, Vamsi Sripathi, Deepthi Karkada, Vivek Menon, Sun Choi, Kushal Datta, and Vikram Saletore. 2019 · 1906
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat. 2019 · 1910
Earlier work this paper cites.
Axial Attention in Multidimensional Transformers
Jonathan Ho, Nal Kalchbrenner, Dirk Weissenborn, and Tim Salimans. 2019 · 1912
Earlier work this paper cites.
Longformer: The Long-Document Transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
Telling BERT’s full story: from Local Attention to Global Aggregation
Damian Pascual, Gino Brunner, and Roger Wattenhofer. 2021 · 2004
Earlier work this paper cites.
Poor Man’s BERT: Smaller and Faster Transformer Models
Hassan Sajjad, Fahim Dalvi, Nadir Durrani, and Preslav Nakov. 2020 · 2004
Earlier work this paper cites.
Hopfield Networks is All You Need
Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl, Michael Widrich, Thomas Adler, Lukas Gruber, Markus Holzleitner, Milena Pavlović, Geir Kjetil Sandve, Victor Greiff, David Kreil, Michael Kopp, Günter Klambauer, Johannes Brandstetter, and Sepp Hochreiter. 2020 · 2008
Earlier work this paper cites.
Efficient Transformers: A Survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2020 · 2009
Earlier work this paper cites.
BinaryBERT: Pushing the Limit of BERT Quantization
Haoli Bai, Wei Zhang, Lu Hou, Lifeng Shang, Jing Jin, Xin Jiang, Qun Liu, Michael Lyu, and Irwin King. 2020 · 2012
Earlier work this paper cites.
Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Cnvlutin: Ineffectual-Neuron-Free Deep Neural Network Computing
J. Albericio, P. Judd, T. Hetherington, T. Aamodt, N. E. Jerger, and A. Moshovos. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
Image Transformer
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran. 2018 · 2018
Cited alongside, same era.
Context-Aware Neural Machine Translation Learns Anaphora Resolution
Elena Voita, Pavel Serdyukov, Rico Sennrich, and Ivan Titov. 2018 · 2018
Cited alongside, same era.
What Does BERT Look at? An Analysis of BERT’s Attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Cited alongside, same era.
Adaptively Sparse Transformers
Gonçalo M. Correia, Vlad Niculae, and André F. T. Martins. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Attention is not not Explanation
Sarah Wiegreffe and Yuval Pinter. 2019 · 2019
Later among the works it cites.
ETC: Encoding Long and Structured Inputs in Transformers
Joshua Ainslie, Santiago Ontanon, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, and Li Yang. 2020 · 2020
Later among the works it cites.
On Identifiability in Transformers
Gino Brunner, Yang Liu, Damian Pascual, Oliver Richter, Massimiliano Ciaramita, and Roger Wattenhofer. 2020 · 2020
Later among the works it cites.
The Lottery Ticket Hypothesis for Pre-trained BERT Networks
Tianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu, Yang Zhang, Zhangyang Wang, and Michael Carbin. 2020 · 2020
Later among the works it cites.
exBERT: A Visual Analysis Tool to Explore Learned Representations in Transformer Models
Benjamin Hoover, Hendrik Strobelt, and Sebastian Gehrmann. 2020 · 2020
Later among the works it cites.
Fully Quantized Transformer for Machine Translation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Attention is not Explanation
Sarthak Jain and Byron C. Wallace. 2019 · 2019
Cited alongside, same era.
Revealing the Dark Secrets of BERT
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 2019
Cited alongside, same era.
Big bidirectional insertion representations for documents
Lala Li and William Chan. 2019 · 2019
Cited alongside, same era.
Are Sixteen Heads Really Better than One?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Cited alongside, same era.
A Multiscale Visualization of Attention in the Transformer Model
Jesse Vig. 2019 · 2019
Cited alongside, same era.
Analyzing the Structure of Attention in a Transformer Language Model
Jesse Vig and Yonatan Belinkov. 2019 · 2019
Cited alongside, same era.
Gabriele Prato, Ella Charlaix, and Mehdi Rezagholizadeh. 2020 · 2020
Later among the works it cites.
A Primer in BERTology: What We Know About How BERT Works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020 · 2020
Later among the works it cites.
Masked Language Model Scoring
Julian Salazar, Davis Liang, Toan Q. Nguyen, and Katrin Kirchhoff. 2020 · 2020
Later among the works it cites.
Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT
Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer. 2020 · 2020
Later among the works it cites.
Similarity Analysis of Contextual Word Representation Models
John Wu, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2020 · 2020
Later among the works it cites.
GOBO: Quantizing Attention-Based NLP Models for Low Latency and Energy Efficient Inference
A. H. Zadeh, I. Edo, O. M. Awad, and A. Moshovos. 2020 · 2020
Later among the works it cites.
TernaryBERT: Distillation-aware Ultra-low Bit BERT
Wei Zhang, Lu Hou, Yichun Yin, Lifeng Shang, Xiao Chen, Xin Jiang, and Qun Liu. 2020 · 2020
Later among the works it cites.
I-BERT: Integer-only BERT Quantization
Sehoon Kim, Amir Gholami, Zhewei Yao, Michael W. Mahoney, and Kurt Keutzer. 2021 · 2021
Closest in time.