Fetching the paper…
Reading the bibliography…
GPT is an auto-regressive Transformer-based pre-trained language model which has attracted a lot of attention in the natural language processing (NLP) domain due to its state-of-the-art performance in several downstream tasks.
Compressing rnns for iot devices by 15-38x using kronecker products
Urmish Thakker, Jesse Beu, Dibakar Gope, Chu Zhou, Igor Fedorov, Ganesh Dasika, and Matthew Mattina. 2019 · 1906
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V Le. 2019 · 1906
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019 · 1909
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019 · 1909
Earlier work this paper cites.
Fully quantized transformer for machine translation
Gabriele Prato, Ella Charlaix, and Mehdi Rezagholizadeh. 2019 · 1910
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
On the history of the kronecker product
Harold V Henderson, Friedrich Pukelsheim, and Shayle R Searle. 1983 · 1983
Earlier work this paper cites.
The ubiquitous kronecker product
Charles F Van Loan. 2000 · 2000
Earlier work this paper cites.
Compressing language models using doped kronecker products
Urmish Thakker, Paul Whatamough, Matthew Mattina, and Jesse Beu. 2020 · 2001
Earlier work this paper cites.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020 · 2002
Earlier work this paper cites.
Mobilebert: a compact task-agnostic bert for resource-limited devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020 · 2004
Cited alongside, same era.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 2005
Cited alongside, same era.
Ternarybert: Distillation-aware ultra-low bit bert
Wei Zhang, Lu Hou, Yichun Yin, Lifeng Shang, Xiao Chen, Xin Jiang, and Qun Liu. 2020 · 2009
Cited alongside, same era.
Kronecker graphs: an approach to modeling networks
Jure Leskovec, Deepayan Chakrabarti, Jon Kleinberg, Christos Faloutsos, and Zoubin Ghahramani. 2010 · 2010
Cited alongside, same era.
Multilevel statistical models , volume 922
Harvey Goldstein. 2011 · 2011
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Openwebtext corpus
Aaron Gokaslan and Vanya Cohen. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Later among the works it cites.
Improving word embedding factorization for compression using distilled nonlinear neural decomposition
Vasileios Lioutas, Ahmad Rashid, Krtin Kumar, Md Akmal Haidar, and Mehdi Rezagholizadeh. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Compressing deep convolutional networks using vector quantization
Yunchao Gong, Liu Liu, Ming Yang, and Lubomir Bourdev. 2014 · 2014
Cited alongside, same era.
Song Han, Huizi Mao, and William J Dally. 2015 · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Cited alongside, same era.
Compression of fully-connected layer in neural network by kronecker product
Shuchang Zhou and Jia-Nan Wu. 2015 · 2015
Cited alongside, same era.
On compressing deep models by low rank and sparse decomposition
Xiyu Yu, Tongliang Liu, Xinchao Wang, and Dacheng Tao. 2017 · 2017
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher
Cited in the paper.
Annealing knowledge distillation
Aref Jafari, Mehdi Rezagholizadeh, Pranav Sharma, and Ali Ghodsi. 2021 · 2021
Closest in time.
Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. 2021 · 2021
Closest in time.
ALP-KD: attention-based layer projection for knowledge distillation
Peyman Passban, Yimeng Wu, Mehdi Rezagholizadeh, and Qun Liu. 2021 · 2021
Closest in time.
Mate-kd: Masked adversarial text, a companion to knowledge distillation
Ahmad Rashid, Vasileios Lioutas, and Mehdi Rezagholizadeh. 2021 · 2021
Closest in time.
Marzieh S. Tahaei, Ella Charlaix, Vahid Partovi Nia, Ali Ghodsi, and Mehdi Rezagholizadeh. 2021 · 2021
Closest in time.