Fetching the paper…
Reading the bibliography…
Pre-trained Transformer-based models have achieved state-of-the-art performance for various Natural Language Processing (NLP) tasks.
Distilling task-specific knowledge from BERT into simple neural networks
Raphael Tang, Yao Lu, Linqing Liu, Lili Mou, Olga Vechtomova, and Jimmy Lin. 2019b · 1903
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b · 1907
Earlier work this paper cites.
Well-read students learn better: The impact of student initialization on knowledge distillation
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 1908
Earlier work this paper cites.
Reweighted proximal pruning for large-scale language representation
Fu-Ming Guo, Sijia Liu, Finlay S Mungall, Xue Lin, and Yanzhi Wang. 2019 · 1909
Earlier work this paper cites.
Megatron-LM: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019 · 1909
Earlier work this paper cites.
Extreme language model compression with optimal subwords and shared projections
Sanqiang Zhao, Raghav Gupta, Yang Song, and Denny Zhou. 2019b · 1909
Earlier work this paper cites.
Distilling BERT into simple neural networks with unlabeled transfer data
Subhabrata Mukherjee and Ahmed H. Awadallah. 2019 · 1910
Earlier work this paper cites.
MKD: A multi-task knowledge distillation approach for pretrained language models
Linqing Liu, Huan Wang, Jimmy Lin, Richard Socher, and Caiming Xiong. 2019a · 1911
Earlier work this paper cites.
WaLDORf: Wasteless language-model distillation on reading-comprehension
James Yi Tian, Alexander P Kreuzer, Pai-Hung Chen, and Hans-Martin Will. 2019 · 1912
Earlier work this paper cites.
Explicit sparse transformer: Concentrated attention through explicit selection
Guangxiang Zhao, Junyang Lin, Zhiyuan Zhang, Xuancheng Ren, Qi Su, and Xu Sun. 2019a · 1912
Earlier work this paper cites.
MiniLM: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020c · 2002
Earlier work this paper cites.
Poor man’s BERT: Smaller and faster transformer models
Hassan Sajjad, Fahim Dalvi, Nadir Durrani, and Preslav Nakov. 2020 · 2004
Earlier work this paper cites.
LightPAFF: A two-stage distillation framework for pre-training and fine-tuning
Kaitao Song, Hao Sun, Xu Tan, Tao Qin, Jianfeng Lu, Hongzhi Liu, and Tie-Yan Liu. 2020 · 2004
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2005
Earlier work this paper cites.
Synthesizer: Rethinking self-attention in transformer models
Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng. 2020 · 2005
Earlier work this paper cites.
Distilling knowledge from pre-trained language models via text smoothing
Xing Wu, Yibing Liu, Xiangyang Zhou, and Dianhai Yu. 2020 · 2005
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Li, Madian Khabsa, Han Fang, and Hao Ma. 2020b · 2006
Earlier work this paper cites.
The lottery ticket hypothesis for pre-trained BERT networks
Tianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu, Yang Zhang, Zhangyang Wang, and Michael Carbin. 2020 · 2007
Earlier work this paper cites.
Finding fast transformers: One-shot neural architecture search by component composition
Henry Tsai, Jayden Ooi, Chun-Sung Ferng, Hyung Won Chung, and Jason Riesa. 2020 · 2008
Earlier work this paper cites.
Weight squeezing: Reparameterization for extreme compression and fast inference
Artem Chumachenko, Daniil Gavrilov, Nikita Balagansky, and Pavel Kalaidin. 2020 · 2010
Earlier work this paper cites.
EdgeBERT: Sentence-level energy optimizations for latency-aware multi-task NLP inference
Thierry Tambe, Coleman Hooper, Lillian Pentecost, Tianyu Jia, En-Yu Yang, Marco Donato, Victor Sanh, Paul Whatmough, Alexander M Rush, David Brooks, et al. 2020 · 2011
Earlier work this paper cites.
Multi-objective optimization
Kalyanmoy Deb. 2014 · 2014
Earlier work this paper cites.
Results of the WMT14 metrics shared task
Matous Machacek and Ondrej Bojar. 2014 · 2014
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Deanand. 2016 · 2016
Earlier work this paper cites.
A survey of model compression and acceleration for deep neural networks
Yu Cheng, Duo Wang, Pan Zhou, and Tao Zhang. 2017 · 2017
Earlier work this paper cites.
Quantized neural networks: Training neural networks with low precision weights and activations
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Transformer to CNN: Label-scarce distillation for efficient text classification
Yew Ken Chia, Sam Witteveen, and Martin Andrews. 2018 · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin. 2018 · 2018
Cited alongside, same era.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Cited alongside, same era.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
PoWER-BERT: Accelerating BERT inference via progressive word-vector elimination
Saurabh Goyal, Anamitra Roy Choudhury, Saurabh Raje, Venkatesan Chakaravarthy, Yogish Sabharwal, and Ashish Verma. 2020 · 2020
Closest in time.
DynaBERT: Dynamic BERT with adaptive width and depth
Lu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang, Xiao Chen, and Qun Liu. 2020 · 2020
Closest in time.
TinyBERT: Distilling BERT for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020 · 2020
Closest in time.
schuBERT: Optimizing elements of BERT
Ashish Khetan and Zohar Karnin. 2020 · 2020
Closest in time.
Efficient transformer-based large scale language representations using hardware-friendly block structured pruning
Bingbing Li, Zhenglun Kong, Tianyun Zhang, Ji Li, Zhengang Li, Hang Liu, and Caiwen Ding. 2020a · 2020
Closest in time.
BERT-EMD: Many-to-many layer mapping for BERT compression with earth mover’s distance
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018 · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Know what you don’t know: Unanswerable questions for SQuAD
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
What does BERT look at? An analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Reducing transformer depth on demand with structured dropout
Angela Fan, Edouard Grave, and Armand Joulin. 2019 · 2019
Cited alongside, same era.
Jianquan Li, Xiaokang Liu, Honghong Zhao, Ruifeng Xu, Min Yang, and Yaohong Jin. 2020b · 2020
Closest in time.
Pruning redundant mappings in transformer models via spectral-normalized identity prior
Zi Lin, Jeremiah Liu, Zi Yang, Nan Hua, and Dan Roth. 2020 · 2020
Closest in time.
FastBERT: a self-distilling BERT with adaptive inference time
Weijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao, Haotang Deng, and Qi Ju. 2020 · 2020
Closest in time.
LadaBERT: Lightweight adaptation of BERT through hybrid model compression
Yihuan Mao, Yujing Wang, Chufan Wu, Chen Zhang, Yang Wang, Quanlu Zhang, Yaming Yang, Yunhai Tong, and Jing Bai. 2020 · 2020
Closest in time.
XtremeDistil: Multi-stage distillation for massive multilingual models
Subhabrata Mukherjee and Ahmed H. Awadallah. 2020 · 2020
Closest in time.
Compressing pre-trained language models by matrix decomposition
Matan Ben Noach and Yoav Goldberg. 2020 · 2020
Closest in time.
Compressing transformer-based semantic parsing models using compositional code embeddings
Prafull Prakash, Saurabh Kumar Shashidhar, Wenlong Zhao, Subendhu Rongali, Haidar Khan, and Michael Kayser. 2020 · 2020
Closest in time.
When BERT plays the lottery, all tickets are winning
Sai Prasanna, Anna Rogers, and Anna Rumshisky. 2020 · 2020
Closest in time.
Pre-trained models for natural language processing: A survey
XiPeng Qiu, TianXiang Sun, YiGe Xu, YunFan Shao, Ning Dai, and XuanJing Huang. 2020 · 2020
Closest in time.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Closest in time.
Fixed Encoder Self-Attention Patterns in Transformer-Based Machine Translation
Alessandro Raganato, Yves Scherrer, and Jörg Tiedemann. 2020 · 2020
Closest in time.
A primer in BERTology: What we know about how BERT works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020 · 2020
Closest in time.
Turing-NLG: A 17-billion-parameter language model by microsoft
Corby Rosset. 2020 · 2020
Closest in time.
Movement pruning: Adaptive sparsity by fine-tuning
Victor Sanh, Thomas Wolf, and Alexander Rush. 2020 · 2020
Closest in time.
Q-BERT: Hessian based ultra low precision quantization of BERT
Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer. 2020 · 2020
Closest in time.
Contrastive distillation on intermediate representations for language model compression
Siqi Sun, Zhe Gan, Yuwei Fang, Yu Cheng, Shuohang Wang, and Jingjing Liu. 2020a · 2020
Closest in time.
Exploring the boundaries of low-resource BERT distillation
Moshe Wasserblat, Oren Pereg, and Peter Izsak. 2020 · 2020
Closest in time.
DeeBERT: Dynamic early exiting for accelerating BERT inference
Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin. 2020 · 2020
Closest in time.
BERT-of-Theseus: Compressing BERT by progressive module replacing
Canwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei, and Ming Zhou. 2020 · 2020
Closest in time.
GOBO: Quantizing attention-based NLP models for low latency and energy efficient inference
Ali H. Zadeh, Isak Edo, Omar Mohamed Awad, and Andreas Moshovos. 2020 · 2020
Closest in time.
Training with quantization noise for extreme model compression
Angela Fan, Pierre Stock, Benjamin Graham, Edouard Grave, Rémi Gribonval, Hervé Jégou, and Armand Joulin. 2021 · 2021
Closest in time.