Fetching the paper…
Reading the bibliography…
This paper presents a novel pre-trained language models (PLM) compression approach based on the matrix product operator (short as MPO) from quantum many-body physics.
Tensorized embedding layers for efficient model compression
Valentin Khrulkov, Oleksii Hrinchuk, Leyla Mirvakhabova, and Ivan Oseledets. 2019 · 1901
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat. 2019 · 1910
Earlier work this paper cites.
The expression of a tensor or a polyadic as a sum of products
Frank L Hitchcock. 1927 · 1927
Earlier work this paper cites.
Some mathematical notes on three-mode factor analysis
Ledyard R Tucker. 1966 · 1966
Earlier work this paper cites.
[8] singular value decomposition: Application to analysis of experimental data
ER Henry and J Hofrichter. 1992 · 1992
Earlier work this paper cites.
Compressing large-scale transformer-based models: A case study on bert
Prakhar Ganesh, Yao Chen, Xin Lou, Mohammad Ali Khan, Yin Yang, Deming Chen, Marianne Winslett, Hassan Sajjad, and Preslav Nakov. 2020 · 2002
Earlier work this paper cites.
Entanglement entropy and quantum field theory
Pasquale Calabrese and John Cardy. 2004 · 2004
Earlier work this paper cites.
Squeezebert: What can computer vision teach nlp about efficient neural networks?
Forrest N Iandola, Albert E Shaw, Ravi Krishna, and Kurt W Keutzer. 2020 · 2006
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Li, Madian Khabsa, Han Fang, and Hao Ma. 2020 · 2006
Earlier work this paper cites.
Matrix product states algorithms and continuous systems
S Iblisdir, R Orus, and JI Latorre. 2007 · 2007
Earlier work this paper cites.
Matrix product operator representations
Bogdan Pirvu, Valentin Murg, J Ignacio Cirac, and Frank Verstraete. 2010 · 2010
Earlier work this paper cites.
Tensor-train decomposition
Ivan V Oseledets. 2011 · 2011
Earlier work this paper cites.
Edgebert: Optimizing on-chip inference for multi-task nlp
Thierry Tambe, Coleman Hooper, Lillian Pentecost, En-Yu Yang, Marco Donato, Victor Sanh, Alexander M Rush, David Brooks, and Gu-Yeon Wei. 2020 · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Tensorizing neural networks
Alexander Novikov, Dmitry Podoprikhin, Anton Osokin, and Dmitry P. Vetrov. 2015 · 2015
Cited alongside, same era.
Ultimate tensorization: compressing convolutional and fc layers alike
Timur Garipov, Dmitry Podoprikhin, Alexander Novikov, and Dmitry Vetrov. 2016 · 2016
Cited alongside, same era.
Quantized neural networks: Training neural networks with low precision weights and activations
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. 2017 · 2017
Cited alongside, same era.
Learning multiple visual domains with residual adapters
Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. 2017 · 2017
Cited alongside, same era.
Compressing recurrent neural network with tensor train
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Tinybert: Distilling BERT for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020 · 2020
Later among the works it cites.
schuBERT: Optimizing elements of BERT
Ashish Khetan and Zohar Karnin. 2020 · 2020
Later among the works it cites.
ALBERT: A lite BERT for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 2020
Later among the works it cites.
Exploring versatile generative language model via parameter-efficient transfer learning
Zhaojiang Lin, Andrea Madotto, and Pascale Fung. 2020 · 2020
Later among the works it cites.
Fastbert: a self-distilling BERT with adaptive inference time
Weijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao, Haotang Deng, and Qi Ju. 2020 · 2020
Later among the works it cites.
Compressing pre-trained language models by matrix decomposition
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Long-term forecasting using tensor-train rnns
Rose Yu, Stephan Zheng, Anima Anandkumar, and Yisong Yue. 2017 · 2017
Cited alongside, same era.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for nlp
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Cited alongside, same era.
A tensorized transformer for language modeling
Xindian Ma, Peng Zhang, Shuai Zhang, Nan Duan, Yuexian Hou, Ming Zhou, and Dawei Song. 2019 · 2019
Cited alongside, same era.
Matan Ben Noach and Yoav Goldberg. 2020 · 2020
Later among the works it cites.
Grounded compositional outputs for adaptive language modeling
Nikolaos Pappas, Phoebe Mulcaire, and Noah A Smith. 2020 · 2020
Later among the works it cites.
How fine can fine-tuning be? learning efficient language models
Evani Radiya-Dixit and Xin Wang. 2020 · 2020
Later among the works it cites.
Contrastive distillation on intermediate representations for language model compression
Siqi Sun, Zhe Gan, Yuwei Fang, Yu Cheng, Shuohang Wang, and Jingjing Liu. 2020a · 2020
Later among the works it cites.
Mobilebert: a compact task-agnostic BERT for resource-limited devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020c · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Julien Chaumond, Lysandre Debut, Victor Sanh, Clement Delangue, Anthony Moi, Pierric Cistac, Morgan Funtowicz, Joe Davison, Sam Shleifer, et al. 2020 · 2020
Later among the works it cites.
Deebert: Dynamic early exiting for accelerating bert inference
Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin. 2020 · 2020
Later among the works it cites.
Bertscore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Later among the works it cites.
Masking as an efficient alternative to finetuning for pretrained language models
Mengjie Zhao, Tao Lin, Fei Mi, Martin Jaggi, and Hinrich Schütze. 2020 · 2020
Later among the works it cites.