Fetching the paper…
Reading the bibliography…
Recently, pre-trained Transformer based language models, such as BERT, have shown great superiority over the traditional methods in many Natural Language Processing (NLP) tasks.
Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat · 1910
Earlier work this paper cites.
Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat · 1910
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Babak Hassibi and David G Stork · 1993
Earlier work this paper cites.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou · 2002
Earlier work this paper cites.
Attentivenas: Improving neural architecture search via attentive sampling
Dilin Wang, Meng Li, Chengyue Gong, and Vikas Chandra · 2011
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Neural networks with few multiplications
Zhouhan Lin, Matthieu Courbariaux, Roland Memisevic, and Yoshua Bengio · 2015
Earlier work this paper cites.
Hardware-oriented approximation of convolutional neural networks
Philipp Gysel, Mohammad Motamedi, and Soheil Ghiasi · 2016
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le · 2016
Earlier work this paper cites.
Quora question pairs
Z. Chen, H. Zhang, X. Zhang, and L. Zhao · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Ristretto: A framework for empirical study of resource-efficient inference in convolutional neural networks
Philipp Gysel, Jon Pimentel, Mohammad Motamedi, and Soheil Ghiasi · 2018
Cited alongside, same era.
Snip: Single-shot network pruning based on connection sensitivity
Namhoon Lee, Thalaiyasingam Ajanthan, and Philip HS Torr · 2018
Cited alongside, same era.
Model compression via distillation and quantization
Antonio Polino, Razvan Pascanu, and Dan Alistarh · 2018
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Fully quantized transformer for improved translation
Gabriele Prato, Ella Charlaix, and Mehdi Rezagholizadeh · 2019
Later among the works it cites.
And the bit goes down: Revisiting the quantization of neural networks
Pierre Stock, Armand Joulin, Rémi Gribonval, Benjamin Graham, and Hervé Jégou · 2019
Later among the works it cites.
EfficientNet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc V Le · 2019
Later among the works it cites.
Training with quantization noise for extreme model compression
Angela Fan, Pierre Stock, Benjamin Graham, Edouard Grave, Rémi Gribonval, Hervé Jégou, and Armand Joulin · 2020
Later among the works it cites.
Wrapnet: Neural net inference with ultra-low-resolution arithmetic
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Cited alongside, same era.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R Bowman · 2018
Cited alongside, same era.
Neural architecture search: A survey
Thomas Elsken, Jan Hendrik Metzen, Frank Hutter, et al · 2019
Cited alongside, same era.
Learned step size quantization
Steven K Esser, Jeffrey L McKinstry, Deepika Bablani, Rathinakumar Appuswamy, and Dharmendra S Modha · 2019
Cited alongside, same era.
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu · 2019
Cited alongside, same era.
Pruning a bert-based question answering model
J Scott McCarley · 2019
Cited alongside, same era.
Variational information distillation for knowledge transfer
Sungsoo Ahn, Shell Xu Hu, Andreas Damianou, Neil D Lawrence, and Zhenwen Dai
Cited in the paper.
The pascal recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini
Cited in the paper.
Renkun Ni, Hong-min Chu, Oscar Castañeda, Ping-yeh Chiang, Christoph Studer, and Tom Goldstein · 2020
Later among the works it cites.
Lookahead: a far-sighted alternative of magnitude-based pruning
Sejun Park, Jaeho Lee, Sangwoo Mo, and Jinwoo Shin · 2020
Later among the works it cites.
Ternarybert: Distillation-aware ultra-low bit bert
Wei Zhang, Lu Hou, Yichun Yin, Lifeng Shang, Xiao Chen, Xin Jiang, and Qun Liu · 2020
Later among the works it cites.
KDLSQ-BERT: A quantized bert combining knowledge distillation with learned step size quantization
Jing Jin, Cai Liang, Tiancheng Wu, Liqin Zou, and Zhiliang Gan · 2021
Later among the works it cites.
Degree-quant: Quantization-aware training for graph neural networks
Shyam A Tailor, Javier Fernandez-Marques, and Nicholas D Lane · 2021
Later among the works it cites.
Hessian-aware pruning and optimal neural implant
Shixing Yu, Zhewei Yao, Amir Gholami, Zhen Dong, Michael W Mahoney, and Kurt Keutzer · 2021
Later among the works it cites.