Fetching the paper…
Reading the bibliography…
The development of over-parameterized pre-trained language models has made a significant contribution toward the success of natural language processing.
Tensorized embedding layers for efficient model compression
Valentin Khrulkov, Oleksii Hrinchuk, Leyla Mirvakhabova, and Ivan Oseledets. 2019 · 1901
Earlier work this paper cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
R Thomas McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 1902
Earlier work this paper cites.
Paws: Paraphrase adversaries from word scrambling
Yuan Zhang, Jason Baldridge, and Luheng He. 2019 · 1904
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V Le. 2019 · 1906
Earlier work this paper cites.
Patient knowledge distillation for bert model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019 · 1908
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019 · 1909
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019 · 1909
Earlier work this paper cites.
Extreme language model compression with optimal subwords and shared projections
Sanqiang Zhao, Raghav Gupta, Yang Song, and Denny Zhou. 2019 · 1909
Earlier work this paper cites.
Fully quantized transformer for machine translation
Gabriele Prato, Ella Charlaix, and Mehdi Rezagholizadeh. 2019 · 1910
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Handbook of matrices
Helmut Lutkepohl. 1997 · 1997
Earlier work this paper cites.
The ubiquitous kronecker product
Charles F Van Loan. 2000 · 2000
Earlier work this paper cites.
Compressing language models using doped kronecker products
Urmish Thakker, Paul Whatamough, Matthew Mattina, and Jesse Beu. 2020 · 2001
Cited alongside, same era.
Pretrained transformers improve out-of-distribution robustness
Dan Hendrycks, Xiaoyuan Liu, Eric Wallace, Adam Dziedzic, Rishabh Krishnan, and Dawn Song. 2020 · 2004
Cited alongside, same era.
Ladabert: Lightweight adaptation of bert through hybrid model compression
Yihuan Mao, Yujing Wang, Chufan Wu, Chen Zhang, Yang Wang, Yaming Yang, Quanlu Zhang, Yunhai Tong, and Jing Bai. 2020 · 2004
Cited alongside, same era.
Mobilebert: a compact task-agnostic bert for resource-limited devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020b · 2004
Cited alongside, same era.
Model compression
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Later among the works it cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Later among the works it cites.
On compressing deep models by low rank and sparse decomposition
Xiyu Yu, Tongliang Liu, Xinchao Wang, and Dacheng Tao. 2017 · 2017
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Kronecker products and matrix calculus with applications
Alexander Graham. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil. 2006 · 2006
Cited alongside, same era.
Contrastive distillation on intermediate representations for language model compression
Siqi Sun, Zhe Gan, Yu Cheng, Yuwei Fang, Shuohang Wang, and Jingjing Liu. 2020a · 2009
Cited alongside, same era.
Fastformers: Highly efficient transformer models for natural language understanding
Young Jin Kim and Hany Hassan Awadalla. 2020 · 2010
Cited alongside, same era.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Cited alongside, same era.
Compressing deep convolutional networks using vector quantization
Yunchao Gong, Liu Liu, Ming Yang, and Lubomir Bourdev. 2014 · 2014
Cited alongside, same era.
Song Han, Huizi Mao, and William J Dally. 2015 · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Cited alongside, same era.
Compression of fully-connected layer in neural network by kronecker product
Shuchang Zhou and Jia-Nan Wu. 2015 · 2015
Cited alongside, same era.
Taku Kudo and John Richardson. 2018 · 2018
Later among the works it cites.
Slim embedding layers for recurrent neural language models
Zhongliang Li, Raymond Kulhanek, Shaojun Wang, Yunxin Zhao, and Shuang Wu. 2018 · 2018
Later among the works it cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018 · 2018
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Improving word embedding factorization for compression using distilled nonlinear neural decomposition
Vasileios Lioutas, Ahmad Rashid, Krtin Kumar, Md Akmal Haidar, and Mehdi Rezagholizadeh. 2020 · 2020
Later among the works it cites.
Bert-of-theseus: Compressing bert by progressive module replacing
Canwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei, and Ming Zhou. 2020 · 2020
Later among the works it cites.
Extremely small bert models from mixed-vocabulary training
Sanqiang Zhao, Raghav Gupta, Yang Song, and Denny Zhou. 2021 · 2021
Closest in time.