Fetching the paper…
Reading the bibliography…
In recent years, the fields of natural language processing (NLP) and information retrieval (IR) have made tremendous progress thanksto deep learning models like Recurrent Neural Networks (RNNs), Gated Recurrent Units (GRUs) and Long Short-Term Memory (LSTMs)networks, and Transformer [120] based models like Bidirectional Encoder Representations from Transformers (BERT) [24], GenerativePre-training Transformer (GPT-2) [94], Multi-task Deep Neural Network (MT-DNN) [73], Extra-Long Network (XLNet) [134], Text-to-text transfer transformer (T5) [95], T-NLG [98] and GShard [63].
Fish transporters and miracle homes: How compositional distributional semantics can help NP parsing. In EMNLP . 1908–1913
Angeliki Lazaridou, Eva Maria Vecchi, and Marco Baroni. 2013 · 1913
Earlier work this paper cites.
Some mathematical notes on three-mode factor analysis
Ledyard R Tucker. 1966 · 1966
Earlier work this paper cites.
Analysis of individual differences in multidimensional scaling via an N-way generalization of “Eckart-Young” decomposition
J Douglas Carroll and Jih-Jie Chang. 1970 · 1970
Earlier work this paper cites.
Least squares quantization in PCM
Stuart Lloyd. 1982 · 1982
Earlier work this paper cites.
Optimal brain damage. In NIPS . 598–605
Yann LeCun, John S Denker, and Sara A Solla. 1990 · 1990
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon. In NIPS . 164–171
Babak Hassibi and David G Stork. 1993 · 1993
Earlier work this paper cites.
Tensor-train decomposition
Ivan V Oseledets. 2011 · 2011
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville. 2013 · 2013
Earlier work this paper cites.
Predicting parameters in deep learning. In NIPS . 2148–2156
Misha Denil, Babak Shakibi, Laurent Dinh, Marc’Aurelio Ranzato, and Nando De Freitas. 2013 · 2013
Earlier work this paper cites.
word2vec
Tomas Mikolov, Kai Chen, Greg Corrado, Jeffrey Dean, L Sutskever, and G Zweig. 2013 · 2013
Earlier work this paper cites.
Peter Huttenlocher (1931–2013)
Christopher A Walsh. 2013 · 2013
Earlier work this paper cites.
Do deep nets really need to be deep?. In NIPS . 2654–2662
Jimmy Ba and Rich Caruana. 2014 · 2014
Earlier work this paper cites.
Compressing deep convolutional networks using vector quantization
Yunchao Gong, Liu Liu, Ming Yang, and Lubomir Bourdev. 2014 · 2014
Earlier work this paper cites.
Reshaping deep neural network for fast decoding by node-pruning. In ICASSP . IEEE, 245–249
Tianxing He, Yuchen Fan, Yanmin Qian, Tian Tan, and Kai Yu. 2014 · 2014
Earlier work this paper cites.
Fixed-point feedforward deep neural network design using weights +1, 0, and -1. In SiPS . IEEE, 1–6
Kyuyeon Hwang and Wonyong Sung. 2014 · 2014
Earlier work this paper cites.
Proximal Newton-type methods for minimizing composite functions
Jason D Lee, Yuekai Sun, and Michael A Saunders. 2014 · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Haşim Sak, Andrew Senior, and Françoise Beaufays. 2014 · 2014
Earlier work this paper cites.
Wojciech Zaremba and Ilya Sutskever. 2014 · 2014
Earlier work this paper cites.
Hippocampal spine head sizes are highly precise
Thomas M Bartol, Cailey Bromer, Justin Kinney, Michael A Chirillo, Jennifer N Bourne, Kristen M Harris, and Terrence J Sejnowski. 2015 · 2015
Earlier work this paper cites.
Strategies for training large vocabulary neural language models
Welin Chen, David Grangier, and Michael Auli. 2015a · 2015
Earlier work this paper cites.
Binaryconnect: Training deep neural networks with binary weights during propagations. In NIPS . 3123–3131
Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David. 2015 · 2015
Earlier work this paper cites.
Sparse overcomplete word vector representations
Manaal Faruqui, Yulia Tsvetkov, Dani Yogatama, Chris Dyer, and Noah Smith. 2015 · 2015
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally. 2015a · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Earlier work this paper cites.
Neural networks with few multiplications
Zhouhan Lin, Matthieu Courbariaux, Roland Memisevic, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Finding function in form: Compositional character models for open vocabulary word representation
Wang Ling, Tiago Luís, Luís Marujo, Ramón Fernandez Astudillo, Silvio Amir, Chris Dyer, Alan W Black, and Isabel Trancoso. 2015 · 2015
Earlier work this paper cites.
Rounding methods for neural networks with low resolution synaptic weights
Lorenz K Muller and Giacomo Indiveri. 2015 · 2015
Earlier work this paper cites.
Auto-sizing neural networks: With applications to n-gram language models
Kenton Murray and David Chiang. 2015 · 2015
Earlier work this paper cites.
Data-free parameter pruning for deep neural networks
Suraj Srinivas and R Venkatesh Babu. 2015 · 2015
Earlier work this paper cites.
Compressing neural language models by sparse word representations
Yunchuan Chen, Lili Mou, Yan Xu, Ge Li, and Zhi Jin. 2016 · 2016
Earlier work this paper cites.
EIE: efficient inference engine on compressed deep neural network
Song Han, Xingyu Liu, Huizi Mao, Jing Pu, Ardavan Pedram, Mark A Horowitz, and William J Dally. 2016a · 2016
Earlier work this paper cites.
DSD: Dense-sparse-dense training for deep neural networks
Song Han, Jeff Pool, Sharan Narang, Huizi Mao, Enhao Gong, Shijian Tang, Erich Elsen, Peter Vajda, Manohar Paluri, John Tran, et al · 2016
Earlier work this paper cites.
Effective quantization methods for recurrent neural networks
Qinyao He, He Wen, Shuchang Zhou, Yuxin Wu, Cong Yao, Xinyu Zhou, and Yuheng Zou. 2016 · 2016
Earlier work this paper cites.
Loss-aware binarization of deep networks
Lu Hou, Quanming Yao, and James T Kwok. 2016 · 2016
Earlier work this paper cites.
Binarized neural networks. In NIPS . 4107–4115
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
Exploring the limits of language modeling
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. 2016 · 2016
Earlier work this paper cites.
Character-aware neural language models. In AAAI
Yoon Kim, Yacine Jernite, David Sontag, and Alexander M Rush. 2016 · 2016
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M Rush. 2016 · 2016
Earlier work this paper cites.
Fengfu Li, Bo Zhang, and Bin Liu. 2016b · 2016
Earlier work this paper cites.
Assessing the ability of LSTMs to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016 · 2016
Earlier work this paper cites.
Learning compact recurrent neural networks. In ICASSP . IEEE, 5960–5964
Zhiyun Lu, Vikas Sindhwani, and Tara N Sainath. 2016 · 2016
Earlier work this paper cites.
Representational distance learning for deep neural networks
Patrick McClure and Nikolaus Kriegeskorte. 2016 · 2016
Earlier work this paper cites.
Recurrent neural networks with limited numerical precision
Joachim Ott, Zhouhan Lin, Ying Zhang, Shih-Chii Liu, and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
Dropneuron: Simplifying the structure of deep neural networks
Wei Pan, Hao Dong, and Yike Guo. 2016 · 2016
Earlier work this paper cites.
On the compression of recurrent neural networks with an application to LVCSR acoustic modeling for embedded speech recognition. In ICASSP . IEEE, 5970–5974
Rohit Prabhavalkar, Ouais Alsharif, Antoine Bruguier, and Lan McGraw. 2016 · 2016
Earlier work this paper cites.
Xnor-net: Imagenet classification using binary convolutional neural networks. In ECCV . Springer, 525–542
Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi. 2016 · 2016
Earlier work this paper cites.
Deep model compression: Distilling knowledge from noisy teachers
Bharat Bhusan Sau and Vineeth N Balasubramanian. 2016 · 2016
Cited alongside, same era.
Compression of neural machine translation models via pruning
Abigail See, Minh-Thang Luong, and Christopher D Manning. 2016 · 2016
Cited alongside, same era.
Chenzhuo Zhu, Song Han, Huizi Mao, and William J Dally. 2016 · 2016
Cited alongside, same era.
A survey of model compression and acceleration for deep neural networks
Yu Cheng, Duo Wang, Pan Zhou, and Tao Zhang. 2017 · 2017
Cited alongside, same era.
Sobolev training for neural networks. In NIPS . 4278–4287
Wojciech M Czarnecki, Simon Osindero, Max Jaderberg, Grzegorz Swirszcz, and Razvan Pascanu. 2017 · 2017
Tensorized Embedding Layers for Efficient Model Compression
Valentin Khrulkov, Oleksii Hrinchuk, Leyla Mirvakhabova, and Ivan Oseledets. 2019 · 2019
Later among the works it cites.
ALBERT: A lite BERT for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 2019
Later among the works it cites.
Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao. 2019a · 2019
Later among the works it cites.
Multi-task deep neural networks for natural language understanding
Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao. 2019b · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Visualizing and understanding neural machine translation. In ACL . 1150–1159
Yanzhuo Ding, Yang Liu, Huanbo Luan, and Maosong Sun. 2017 · 2017
Cited alongside, same era.
Ensemble distillation for neural machine translation
Markus Freitag, Yaser Al-Onaizan, and Baskaran Sankaran. 2017 · 2017
Cited alongside, same era.
Network sketching: Exploiting binary structure in deep cnns. In CVPR . 5955–5963
Yiwen Guo, Anbang Yao, Hao Zhao, and Yurong Chen. 2017 · 2017
Cited alongside, same era.
Quantized neural networks: Training neural networks with low precision weights and activations
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. 2017 · 2017
Cited alongside, same era.
Low precision RNNs: Quantizing RNNs without losing accuracy
Supriya Kapur, Asit Mishra, and Debbie Marr. 2017 · 2017
Cited alongside, same era.
Apprentice: Using knowledge distillation techniques to improve low-precision network accuracy
Asit Mishra and Debbie Marr. 2017 · 2017
Cited alongside, same era.
Exploring sparsity in recurrent neural networks
Sharan Narang, Erich Elsen, Gregory Diamos, and Shubho Sengupta. 2017a · 2017
Cited alongside, same era.
Xindian Ma, Peng Zhang, Shuai Zhang, Nan Duan, Yuexian Hou, Dawei Song, and Ming Zhou. 2019 · 2019
Later among the works it cites.
Are Sixteen Heads Really Better than One?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Later among the works it cites.
Improved Knowledge Distillation via Teacher Assistant
Seyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine, Akihiro Matsukawa, and Hassan Ghasemzadeh. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019 · 2019
Later among the works it cites.
Turing-nlg: A 17-billion-parameter language model by microsoft
C Rosset. 2019 · 2019
Later among the works it cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Later among the works it cites.
Q-bert: Hessian based ultra low precision quantization of bert
Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer. 2019 · 2019
Later among the works it cites.
Megatron-lm: Training multi-billion parameter language models using gpu model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019 · 2019
Later among the works it cites.
Energy and policy considerations for deep learning in NLP
Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019 · 2019
Later among the works it cites.
Patient knowledge distillation for bert model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019 · 2019
Later among the works it cites.
Multilingual neural machine translation with knowledge distillation
Xu Tan, Yi Ren, Di He, Tao Qin, Zhou Zhao, and Tie-Yan Liu. 2019 · 2019
Later among the works it cites.
Distilling Task-Specific Knowledge from BERT into Simple Neural Networks
Raphael Tang, Yao Lu, Linqing Liu, Lili Mou, Olga Vechtomova, and Jimmy Lin. 2019 · 2019
Later among the works it cites.
Lightweight and Efficient Neural Natural Language Processing with Quaternion Networks
Yi Tay, Aston Zhang, Luu Anh Tuan, Jinfeng Rao, Shuai Zhang, Shuohang Wang, Jie Fu, and Siu Cheung Hui. 2019 · 2019
Later among the works it cites.
Well-read students learn better: The impact of student initialization on knowledge distillation
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
WEST: Word Encoded Sequence Transducers. In ICASSP . IEEE, 7340–7344
Ehsan Variani, Ananda Theertha Suresh, and Mitchel Weintraub. 2019 · 2019
Later among the works it cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Later among the works it cites.
SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019a · 2019
Later among the works it cites.
Structured Pruning of Large Language Models
Ziheng Wang, Jeremy Wohlwend, and Tao Lei. 2019c · 2019
Later among the works it cites.
XLNet: Generalized Autoregressive Pretraining for Language Understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V Le. 2019 · 2019
Later among the works it cites.
Enhanced bayesian compression via deep reinforcement learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 6946–6955
Xin Yuan, Liangliang Ren, Jiwen Lu, and Jie Zhou. 2019 · 2019
Later among the works it cites.
Extreme Language Model Compression with Optimal Subwords and Shared Projections
Sanqiang Zhao, Raghav Gupta, Yang Song, and Denny Zhou. 2019 · 2019
Later among the works it cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan. 2020 · 2020
Closest in time.
Optimized Transformer Models for FAQ Answering. In PAKDD . To appear
Sonam Damani, Kedhar Nath Narahari, Ankush Chatterjee, Manish Gupta, and Puneet Agrawal. 2020 · 2020
Closest in time.
Model Compression and Hardware Acceleration for Neural Networks: A Comprehensive Survey
Lei Deng, Guoqi Li, Song Han, Luping Shi, and Yuan Xie. 2020 · 2020
Closest in time.
SqueezeBERT: What can computer vision teach NLP about efficient neural networks?
Forrest N Iandola, Albert E Shaw, Ravi Krishna, and Kurt W Keutzer. 2020 · 2020
Closest in time.
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. 2020 · 2020
Closest in time.
Fastformers: Highly efficient transformer models for natural language understanding
Young Jin Kim and Hany Hassan Awadalla. 2020 · 2020
Closest in time.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya. 2020 · 2020
Closest in time.
Gshard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2020 · 2020
Closest in time.
BERT-EMD: Many-to-Many Layer Mapping for BERT Compression with Earth Mover’s Distance
Jianquan Li, Xiaokang Liu, Honghong Zhao, Ruifeng Xu, Min Yang, and Yaohong Jin. 2020 · 2020
Closest in time.
XtremeDistil: Multi-stage Distillation for Massive Multilingual Models. In ACL . 2221–2234
Subhabrata Mukherjee and Ahmed Hassan Awadallah. 2020 · 2020
Closest in time.
When BERT Plays the Lottery, All Tickets Are Winning
Sai Prasanna, Anna Rogers, and Anna Rumshisky. 2020 · 2020
Closest in time.
Blockwise Self-Attention for Long Document Understanding. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Findings . 2555–2565
Jiezhong Qiu, Hao Ma, Omer Levy, Wen-tau Yih, Sinong Wang, and Jie Tang. 2020 · 2020
Closest in time.
Mobilebert: a compact task-agnostic bert for resource-limited devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020 · 2020
Closest in time.
Yi Tay, Dara Bahri, Liu Yang, Donald Metzler, and Da-Cheng Juan. 2020 · 2020
Closest in time.
Knowledge distillation and student-teacher learning for visual intelligence: A review and new outlooks. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2020 . IEEE CVPR
Lin Wang and Kuk-Jin Yoon. 2020 · 2020
Closest in time.
Linformer: Self-Attention with Linear Complexity
Sinong Wang, Belinda Li, Madian Khabsa, Han Fang, and Hao Ma. 2020a · 2020
Closest in time.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020b · 2020
Closest in time.
Bert-of-theseus: Compressing bert by progressive module replacing
Canwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei, and Ming Zhou. 2020 · 2020
Closest in time.
Persistent rnns: Stashing recurrent weights on-chip. In ICML . 2024–2033
Greg Diamos, Shubho Sengupta, Bryan Catanzaro, Mike Chrzanowski, Adam Coates, Erich Elsen, Jesse Engel, Awni Hannun, and Sanjeev Satheesh. 2016 · 2033
Closest in time.
Learning Compact Neural Word Embeddings by Parameter Space Sharing.. In IJCAI . 2046–2052
Jun Suzuki and Masaaki Nagata. 2016 · 2052
Closest in time.