Fetching the paper…
Reading the bibliography…
We provide a new LLM-compression solution via SVD, unlocking new possibilities for LLM compression beyond quantization and pruning.
The singular value decomposition: Its computation and some applications
Virginia Klema and Alan Laub · 1980
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitch Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz · 1993
Earlier work this paper cites.
Candid covariance-free incremental principal component analysis
Juyang Weng, Yilu Zhang, and Wey-Shiuan Hwang · 2003
Earlier work this paper cites.
Mimo transmission over a time-varying channel using svd
Guillaume Lebrun, Jason Gao, and Mike Faulkner · 2005
Earlier work this paper cites.
Model compression
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Image compression using svd
HS Prasantha, HL Shashidhara, and KN Balasubramanya Murthy · 2007
Earlier work this paper cites.
Compression of facial images using the k-svd algorithm
Ori Bryt and Michael Elad · 2008
Earlier work this paper cites.
Low-rank matrix factorization for deep neural network training with high-dimensional output targets
Tara N Sainath, Brian Kingsbury, Vikas Sindhwani, Ebru Arisoy, and Bhuvana Ramabhadran · 2013
Earlier work this paper cites.
Exploiting linear structure within convolutional networks for efficient evaluation
Emily L Denton, Wojciech Zaremba, Joan Bruna, Yann LeCun, and Rob Fergus · 2014
Earlier work this paper cites.
Speeding up convolutional neural networks with low rank expansions
Max Jaderberg, Andrea Vedaldi, and Andrew Zisserman · 2014
Earlier work this paper cites.
Speeding-up convolutional neural networks using fine-tuned cp-decomposition
Vadim Lebedev, Yaroslav Ganin, Maksim Rakhuba, Ivan Oseledets, and Victor Lempitsky · 2014
Earlier work this paper cites.
An efficient svd-based method for image denoising
Qiang Guo, Caiming Zhang, Yunfeng Zhang, and Hui Liu · 2015
Earlier work this paper cites.
Acdc: A structured efficient linear layer
Marcin Moczulski, Misha Denil, Jeremy Appleyard, and Nando de Freitas · 2015
Earlier work this paper cites.
Accelerating very deep convolutional networks for classification and detection
Xiangyu Zhang, Jianhua Zou, Kaiming He, and Jian Sun · 2015
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
P Rajpurkar · 2016
Earlier work this paper cites.
A survey of model compression and acceleration for deep neural networks
Yu Cheng, Duo Wang, Pan Zhou, and Tao Zhang · 2017
Earlier work this paper cites.
A novel strategy for signal denoising using reweighted svd and its applications to weak fault feature enhancement of rotating machinery
Ming Zhao and Xiaodong Jia · 2017
Earlier work this paper cites.
Polycystic kidney disease
Carsten Bergmann, Lisa M Guay-Woodford, Peter C Harris, Shigeo Horie, Dorien JM Peters, and Vicente E Torres · 2018
Cited alongside, same era.
Groupreduce: Block-wise low-rank approximation for neural language model shrinking
Patrick Chen, Si Si, Yang Li, Ciprian Chelba, and Cho-Jui Hsieh · 2018
Cited alongside, same era.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Cited alongside, same era.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Cited alongside, same era.
Online embedding compression for text classification using low rank matrix factorization
Anish Acharya, Rahul Goel, Angeliki Metallinou, and Inderjit Dhillon · 2019
Matrix compression via randomized low rank and low precision factorization
Rajarshi Saha, Varun Srivastava, and Mert Pilanci · 2023
Later among the works it cites.
The truth is in there: Improving reasoning in language models with layer-selective rank reduction
Pratyusha Sharma, Jordan T Ash, and Dipendra Misra · 2023
Later among the works it cites.
A simple and effective pruning approach for large language models
Mingjie Sun, Zhuang Liu, Anna Bair, and J. Zico Kolter · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Sheared llama: Accelerating language model pre-training via structured pruning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mathqa: Towards interpretable math word problem solving with operation-based formalisms
Aida Amini, Saadia Gabriel, Peter Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi · 2019
Cited alongside, same era.
Hellaswag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Cited alongside, same era.
Piqa: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al · 2020
Cited alongside, same era.
Adabert: Task-adaptive bert compression with differentiable neural architecture search
Daoyuan Chen, Yaliang Li, Minghui Qiu, Zhen Wang, Bofang Li, Bolin Ding, Hongbo Deng, Jun Huang, Wei Lin, and Jingren Zhou · 2020
Cited alongside, same era.
A comprehensive survey on model compression and acceleration
Tejalal Choudhary, Vipul Mishra, Anurag Goswami, and Jagannathan Sarangapani · 2020
Cited alongside, same era.
Dynabert: Dynamic bert with adaptive width and depth
Lu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang, Xiao Chen, and Qun Liu · 2020
Cited alongside, same era.
A framework for few-shot language model evaluation
Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, et al · 2021
Cited alongside, same era.
Mengzhou Xia, Tianyu Gao, Zhiyuan Zeng, and Danqi Chen · 2023
Later among the works it cites.
Asvd: Activation-aware singular value decomposition for compressing large language models
Zhihang Yuan, Yuzhang Shang, Yue Song, Qiang Wu, Yan Yan, and Guangyu Sun · 2023
Later among the works it cites.
Slicegpt: Compress large language models by deleting rows and columns
Saleh Ashkboos, Maximilian L Croci, Marcelo Gennari do Nascimento, Torsten Hoefler, and James Hensman · 2024
Later among the works it cites.
Everybody prune now: Structured pruning of llms with only forward passes
Lucio Dery, Steven Kolawole, Jean-François Kagy, Virginia Smith, Graham Neubig, and Ameet Talwalkar · 2024
Later among the works it cites.
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer · 2024
Later among the works it cites.
Openvla: An open-source vision-language-action model
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn · 2024
Later among the works it cites.
Jung Hyun Lee, Jeonghoon Kim, June Yong Yang, Se Jung Kwon, Eunho Yang, Kang Min Yoo, and Dongsoo Lee · 2024
Later among the works it cites.
Awq: Activation-aware weight quantization for on-device llm compression and acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han · 2024
Later among the works it cites.
Improved baselines with visual instruction tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee · 2024
Later among the works it cites.
Compressing large language models using low rank and low precision decomposition
Rajarshi Saha, Naomi Sagan, Varun Srivastava, Andrea J Goldsmith, and Mert Pilanci · 2024
Later among the works it cites.
Svd-llm: Truncation-aware singular value decomposition for large language model compression
Xin Wang, Yu Zheng, Zhongwei Wan, and Mi Zhang · 2024
Later among the works it cites.
The devil is in the details — Wikipedia, the free encyclopedia
Wikipedia · 2024
Later among the works it cites.
Lqer: Low-rank quantization error reconstruction for llms
Cheng Zhang, Jianyi Cheng, George A Constantinides, and Yiren Zhao · 2024
Later among the works it cites.