Fetching the paper…
Reading the bibliography…
We present STAT: a simple algorithm to prune transformer models without any fine-tuning.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 1910
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 1910
Earlier work this paper cites.
Linear least squares solutions by householder transformations
Peter Businger and Gene H Golub · 1965
Earlier work this paper cites.
Some applications of the rank revealing qr factorization
Tony F Chan and Per Christian Hansen · 1992
Earlier work this paper cites.
Rank-revealing qr factorizations and the singular value decomposition
Yoo Pyo Hong and C-T Pan · 1992
Earlier work this paper cites.
On rank-revealing factorisations
Shivkumar Chandrasekaran and Ilse CF Ipsen · 1994
Earlier work this paper cites.
Efficient algorithms for computing a strong rank-revealing QR factorization
M. Gu and S. Eisenstat · 1996
Earlier work this paper cites.
A theory of pseudoskeleton approximations
Sergei A Goreinov, Eugene E Tyrtyshnikov, and Nickolai L Zamarashkin · 1997
Earlier work this paper cites.
A blas-3 version of the qr factorization with column pivoting
Gregorio Quintana-Ortí, Xiaobai Sun, and Christian H. Bischof · 1998
Earlier work this paper cites.
LAPACK Users’ Guide
E. Anderson, Z. Bai, C. Bischof, S. Blackford, J. Demmel, J. Dongarra, J. Du Croz, A. Greenbaum, S. Hammarling, A. McKenney, and D. Sorensen · 1999
Earlier work this paper cites.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou · 2002
Earlier work this paper cites.
A fast direct solver for boundary integral equations in two dimensions
Per-Gunnar Martinsson and V. Rokhlin · 2004
Earlier work this paper cites.
On the compression of low rank matrices
H. Cheng, Z. Gimbutas, Per-Gunnar Martinsson, and V. Rokhlin · 2005
Earlier work this paper cites.
Randomized algorithms for the low-rank approximation of matrices
Edo Liberty, Franco Woolfe, Per-Gunnar Martinsson, Vladimir Rokhlin, and Mark Tygert · 2007
Earlier work this paper cites.
An improved approximation algorithm for the column subset selection problem
Christos Boutsidis, Michael W Mahoney, and Petros Drineas · 2009
Earlier work this paper cites.
On selecting a maximum volume sub-matrix of a matrix and related problems
Ali Civril and Malik Magdon-Ismail · 2009
Earlier work this paper cites.
CUR matrix decompositions for improved data analysis
Michael W Mahoney and Petros Drineas · 2009
Earlier work this paper cites.
Column subset selection, matrix factorization, and eigenvalue optimization
Joel A Tropp · 2009
Earlier work this paper cites.
Finding structure with randomness: Probabilistic algorithms for constructing approximate matrix decompositions
Nathan Halko, Per-Gunnar Martinsson, and Joel A Tropp · 2011
Earlier work this paper cites.
A randomized algorithm for the decomposition of matrices
Per-Gunnar Martinsson, Vladimir Rokhlin, and Mark Tygert · 2011
Earlier work this paper cites.
A fast direct solver for structured linear systems by recursive skeletonization
K.L. Ho and L. Greengard · 2012
Earlier work this paper cites.
Communication avoiding rank revealing qr factorization with column pivoting
James W Demmel, Laura Grigori, Ming Gu, and Hua Xiang · 2015
Earlier work this paper cites.
Hierarchical interpolative factorization for elliptic operators: Integral equations
K.L. Ho and L. Ying · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Data-free parameter pruning for deep neural networks
Suraj Srinivas and R. Venkatesh Babu · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Earlier work this paper cites.
Hierarchical interpolative factorization for elliptic operators: differential equations
Kenneth L Ho and Lexing Ying · 2016
Earlier work this paper cites.
Low-rank approximation and regression in input sparsity time
Kenneth L Clarkson and David P Woodruff · 2017
Cited alongside, same era.
Channel pruning for accelerating very deep neural networks
Yihui He, Xiangyu Zhang, and Jian Sun · 2017
Cited alongside, same era.
A recursive skeletonization factorization based on strong admissibility
Victor Minden, Kenneth L Ho, Anil Damle, and Lexing Ying · 2017
Cited alongside, same era.
Efficient algorithms for cur and interpolative matrix decompositions
Sergey Voronin and Per-Gunnar Martinsson · 2017
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Reducing transformer depth on demand with structured dropout
I-bert: Integer-only bert quantization
Sehoon Kim, Amir Gholami, Zhewei Yao, Michael W. Mahoney, and Kurt Keutzer · 2021
Later among the works it cites.
Block pruning for faster transformers
François Lagunas, Ella Charlaix, Victor Sanh, and Alexander M. Rush · 2021
Later among the works it cites.
Post-training deep neural network pruning via layer-wise calibration
Ivan Lazarevich, Alexander Kozlov, and Nikita Malinin · 2021
Later among the works it cites.
EBERT: Efficient BERT inference with dynamic structured pruning
Zejian Liu, Fanrong Li, Gang Li, and Jian Cheng · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Angela Fan, Edouard Grave, and Armand Joulin · 2019
Cited alongside, same era.
The state of sparsity in deep neural networks
Trevor Gale, Erich Elsen, and Sara Hooker · 2019
Cited alongside, same era.
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu · 2019
Cited alongside, same era.
Fast Direct Solvers for Elliptic PDEs
Per-Gunnar. Martinsson · 2019
Cited alongside, same era.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig · 2019
Cited alongside, same era.
Patient knowledge distillation for bert model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu · 2019
Cited alongside, same era.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov · 2019
Cited alongside, same era.
Archit Parnami, Rahul Singh, and Tarun Joshi · 2021
Later among the works it cites.
Layer-wise pruning of transformer attention heads for efficient language modeling
Kyuhong Shim, Iksoo Choi, Wonyong Sung, and Jungwook Choi · 2021
Later among the works it cites.
Does knowledge distillation really work?
Samuel Stanton, Pavel Izmailov, Polina Kirichenko, A. Alemi Alexander, and Andrew Gordon Wilson · 2021
Later among the works it cites.
No fine-tuning, no cry: Robust SVD for compressing deep networks
Murad Tukan, Alaa Maalouf, Matan Weksler, and Dan Feldman · 2021
Later among the works it cites.
Red++ : Data-free pruning of deep neural networks via input splitting and output merging
Edouard Yvinec, Arnaud Dapogny, Matthieu Cord, and Kevin Bailly · 2021
Later among the works it cites.
Model preserving compression for neural networks
Jerry Chee, Megan Flynn (née Renz), Anil Damle, and Christopher M De Sa · 2022
Later among the works it cites.
Spdy: Accurate pruning with speedup guarantees
Elias Frantar and Dan Alistarh · 2022
Later among the works it cites.
The optimal bert surgeon: Scalable and accurate second-order pruning for large language models
Eldar Kurtic, Daniel Campos, Tuan Nguyen, Elias Frantar, Mark Kurtz, Benjamin Fineran, Michael Goin, and Dan Alistarh · 2022
Later among the works it cites.
A fast post-training pruning framework for transformers
Wooksuk Kwon, Sehoon Kim, Michael Mahoney, Joseph Hassoun, Kurt Keutzer, and Amir Gholami · 2022
Later among the works it cites.
On the effect of dropping layers of pre-trained transformer models
Hassan Sajjad, Fahim Dalvi, Nadir Durrani, and Preslav Nakov · 2022
Later among the works it cites.
Structured pruning learns compact and accurate models
Mengzhou Xia, Zexuan Zhong, and Danqi Chen · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models, 2022
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer · 2022
Later among the works it cites.
Structure-aware analyses and algorithms for interpolative decompositions
Robin Armstrong, Alex Buzali, and Anil Damle · 2023
Later among the works it cites.
QuIP: 2-bit quantization of large language models with guarantees
Jerry Chee, Yaohui Cai, Volodymyr Kuleshov, and Christopher De Sa · 2023
Later among the works it cites.
Simpler is better: a comparative study of randomized pivoting algorithms for cur and interpolative decompositions
Yijun Dong and Per-Gunnar Martinsson · 2023
Later among the works it cites.
Sparsegpt: Massive language models can be accurately pruned in one-shot
Elias Frantar and Dan Alistarh · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom · 2023
Later among the works it cites.
SliceGPT: Compress large language models by deleting rows and columns
Saleh Ashkboos, Maximilian L. Croci, Marcelo Gennari do Nascimento, Torsten Hoefler, and James Hensman · 2024
Closest in time.
Extreme compression of large language models via additive quantization, 2024
Vage Egiazarian, Andrei Panferov, Denis Kuznedelev, Elias Frantar, Artem Babenko, and Dan Alistarh · 2024
Closest in time.
Omniquant: Omnidirectionally calibrated quantization for large language models
Wenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu, Lirui Zhao, Zhiqian Li, Kaipeng Zhang, Peng Gao, Yu Qiao, and Ping Luo · 2024
Closest in time.
A simple and effective pruning approach for large language models
Mingjie Sun, Zhuang Liu, Anna Bair, and J Zico Kolter · 2024
Closest in time.
The LLM surgeon
Tycho F. A. van der Ouderaa, Markus Nagel, Mart Van Baalen, and Tijmen Blankevoort · 2024
Closest in time.