Fetching the paper…
Reading the bibliography…
This paper focuses on modern efficient training and inference technologies on foundation models and illustrates them from two perspectives: model and system design.
Optimal brain damage
Yann LeCun, John Denker, and Sara Solla · 1989
Earlier work this paper cites.
Adaptive mixtures of local experts
Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan, and Geoffrey E. Hinton · 1991
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Babak Hassibi and David Stork · 1992
Earlier work this paper cites.
Hierarchical mixtures of experts and the em algorithm
Michael I. Jordan and Robert A. Jacobs · 1994
Earlier work this paper cites.
A parallel mixture of svms for very large scale problems
Ronan Collobert, Samy Bengio, and Yoshua Bengio · 2001
Earlier work this paper cites.
Infinite mixtures of gaussian process experts
Carl Rasmussen and Zoubin Ghahramani · 2001
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding, 2020a
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2006
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding, 2020b
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2006
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Learning factored representations in a deep mixture of experts, 2014
David Eigen, Marc’Aurelio Ranzato, and Ilya Sutskever · 2014
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Song Han, Jeff Pool, John Tran, and William Dally · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio · 2015
Earlier work this paper cites.
Generative image modeling using spatial lstms
Lucas Theis and Matthias Bethge · 2015
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M Rush · 2016
Earlier work this paper cites.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients
Shuchang Zhou, Yuxin Wu, Zekun Ni, Xinyu Zhou, He Wen, and Yuheng Zou · 2016
Earlier work this paper cites.
Expert gate: Lifelong learning with a network of experts
Rahaf Aljundi, Punarjay Chakravarty, and Tinne Tuytelaars · 2017
Earlier work this paper cites.
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer, 2017
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
Sergey Zagoruyko and Nikos Komodakis · 2017
Earlier work this paper cites.
Multimodal machine learning: A survey and taxonomy
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency · 2018
Earlier work this paper cites.
Mixed precision training of convolutional neural networks using integer operations
Dipankar Das, Naveen Mellempudi, Dheevatsa Mudigere, Dhiraj Kalamkar, Sasikanth Avancha, Kunal Banerjee, Srinivas Sridharan, Karthik Vaidyanathan, Bharat Kaul, Evangelos Georganas, Alexander Heinecke, Pradeep Dubey, Jesus Corbal, Nikita Shustrov, Roma Dubtsov, Evarist Fomenko, and Vadim Pirogov · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2018
Earlier work this paper cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko · 2018
Earlier work this paper cites.
Efficient 8-bit quantization of transformer neural machine language translation model
Aishwarya Bhandare, Vamsi Sripathi, Deepthi Karkada, Vivek Menon, Sun Choi, Kushal Datta, and Vikram Saletore · 2019
Earlier work this paper cites.
Sparse networks from scratch: Faster training without losing performance
Tim Dettmers and Luke Zettlemoyer · 2019
Earlier work this paper cites.
Learned step size quantization
Steven K. Esser, Jeffrey L. McKinstry, Devesh Bablani, Rathinakumar Appuswamy, and Dharmendra S. Modha · 2019
Earlier work this paper cites.
The state of sparsity in deep neural networks
Trevor Gale, Erich Elsen, and Sara Hooker · 2019
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mingxing Chen, Zhenzhong Chen, Fei Tan, Yashuo Sugawara, Wei Wang, Maxim Krikun, et al · 2019
Earlier work this paper cites.
Distilling knowledge via knowledge transfer in text-to-text models
Xinzhe Chen, Yuan Gao, Jinsong Zhang, and Xiaodong Zhou · 2020
Earlier work this paper cites.
Model compression and hardware acceleration for neural networks: A comprehensive survey
Lei Deng, Guoqi Li, Song Han, Luping Shi, and Yuan Xie · 2020
Earlier work this paper cites.
Dynabert: Dynamic bert with adaptive width and depth
Le Hou, Zhewei Yu, Fei Chen, Peng Jin, Zhenyu Yang, Yen-Kuang Cheng, and Lin Xu · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and Francois Fleuret · 2020
Earlier work this paper cites.
Improved knowledge distillation via teacher assistant
Seyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Niranjan Levine, Akihiro Matsukawa, and Hassan Ghasemzadeh · 2020
Earlier work this paper cites.
Mixed-precision deep learning based on computational memory
S. R. Nandakumar, Manuel Le Gallo, Christophe Piveteau, Vinay Joshi, Giovanni Mariani, Irem Boybat, Geethan Karunaratne, Riduan Khaddam-Aljameh, Urs Egger, Anastasios Petropoulos, Theodore Antonakopoulos, Bipin Rajendran, Abu Sebastian, and Evangelos Eleftheriou · 2020
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020
Earlier work this paper cites.
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamás Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al · 2021
Earlier work this paper cites.
Reasoning with transformer-based models: Deep learning, but shallow reasoning
Chadi Helwe, Chloé Clavel, and Fabian M. Suchanek · 2021
Earlier work this paper cites.
Perceiver io: A general architecture for structured inputs & outputs
Andrew Jaegle et al · 2021
Earlier work this paper cites.
Sequence-level knowledge distillation for low-resource neural machine translation
Vahid Jafari, Haitao Liu, and Qing Lin · 2021
Earlier work this paper cites.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang · 2021
Earlier work this paper cites.
Accelerating sparse deep neural networks
Asit Mishra, Jorge Albericio Latorre, Jeff Pool, Darko Stosic, Dusan Stosic, Ganesh Venkatesh, Chong Yu, and Paulius Micikevicius · 2021
Earlier work this paper cites.
A white paper on neural network quantization
Markus Nagel, Marios Fournarakis, Rana Ali Amjad, Yelysei Bondarenko, Mart van Baalen, and Tijmen Blankevoort · 2021
Earlier work this paper cites.
Fastertransformer: An efficient transformer library
NVIDIA · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Earlier work this paper cites.
Zero-infinity: breaking the gpu memory wall for extreme scale deep learning
Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, and Yuxiong He · 2021
Earlier work this paper cites.
Dynamicvit: Efficient vision transformers with dynamic token sparsification
Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh · 2021
Earlier work this paper cites.
ZeRO-Offload: Democratizing Billion-Scale model training
Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He · 2021
Earlier work this paper cites.
Scaling vision with sparse mixture of experts
Carlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann, Rodolphe Jenatton, André Susano Pinto, Daniel Keysers, and Neil Houlsby · 2021
Earlier work this paper cites.
Automatic mixed-precision quantization search of bert
Changsheng Zhao, Ting Hua, Yilin Shen, Qian Lou, and Hongxia Jin · 2021
Earlier work this paper cites.
Flamingo: A visual language model for few-shot learning
Jean-Baptiste Alayrac et al · 2022
Earlier work this paper cites.
Efficient large scale language modeling with mixtures of experts, 2022
Mikel Artetxe and Shruti Bhosale · 2022
Earlier work this paper cites.
Metro: Efficient denoising pretraining of large scale autoencoding language models with model generated signals (arxiv: 2204.06644). arxiv
P Bajaj, C Xiong, G Ke, X Liu, D He, S Tiwary, TY Liu, P Bennett, X Song, and J Gao · 2022
Cited alongside, same era.
Ta-moe: Topology-aware large scale mixture-of-expert training
Chang Chen, Min Li, Zhihua Wu, Dianhai Yu, and Chao Yang · 2022
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Tri Dao et al · 2022
Cited alongside, same era.
GLaM: Efficient scaling of language models with mixture-of-experts
Nan Du and Huang · 2022
Cited alongside, same era.
Optimal brain compression: A framework for accurate post-training quantization and pruning
Elias Frantar and Dan Alistarh · 2022
Cited alongside, same era.
Smoothquant: Accurate and efficient post-training quantization for large language models
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han · 2023
Later among the works it cites.
Efficientdm: Robust and efficient quantization for diffusion models
Huanrui Yang, Zhenyu Zhang, Yiran Chen, and Hai Li · 2023
Later among the works it cites.
Quant-llm: Post-training quantization for large language models
Zhewei Yao, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer · 2023
Later among the works it cites.
Edgemoe: Fast on-device inference of moe-based large language models, 2023
Rongjie Yi, Liwei Guo, Shiyun Wei, Ao Zhou, Shangguang Wang, and Mengwei Xu · 2023
Later among the works it cites.
Rptq: Reorder-based post-training quantization for large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xinyang Geng, Hao Liu, Lisa Lee, Dale Schuurmans, Sergey Levine, and Pieter Abbeel · 2022
Cited alongside, same era.
A survey of quantization methods for efficient neural network inference
Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W Mahoney, and Kurt Keutzer · 2022
Cited alongside, same era.
Preserving knowledge in pre-trained language models via knowledge distillation
Yeyun Gong, Jinglu Chen, Changjian Wang, Han Zhang, Haoyang Liu, Yu Sun, Weizhu Chen, Bowen Zhou, and Tie-Yan Liu · 2022
Cited alongside, same era.
Entropy regularization methods for parameter space exploration
Shuai Han, Wenbo Zhou, Shuai Lü, Sheng Zhu, and Xiaoyu Gong · 2022
Cited alongside, same era.
Fastermoe: modeling and optimizing training of large-scale dynamic pre-trained models
Jiaao He, Jidong Zhai, Tiago Antunes, Haojie Wang, Fuwen Luo, Shangfeng Shi, and Qin Li · 2022
Cited alongside, same era.
Causal linear transformers: Overcoming cumulative sum bottlenecks in autoregressive inference
Wen Hua, Zhen Qin, Dong Li, Weigao Sun, and Yiran Zhong · 2022
Cited alongside, same era.
F8net: Fixed-point 8-bit only multiplication for network quantization
Qing Jin, Jian Ren, Richard Zhuang, Sumant Hanumante, Zhengang Li, Zhiyu Chen, Yanzhi Wang, Kaiyuan Yang, and Sergey Tulyakov · 2022
Cited alongside, same era.
Zhihang Yuan, Lin Niu, Jiawei Liu, Wenyu Liu, Xinggang Wang, Yuzhang Shang, Guangyu Sun, Qiang Wu, Jiaxiang Wu, and Bingzhe Wu · 2023
Later among the works it cites.
SmartMoE: Efficiently training Sparsely-Activated models through combining offline and online parallelization
Mingshu Zhai, Jiaao He, Zixuan Ma, Zan Zong, Runqing Zhang, and Jidong Zhai · 2023
Later among the works it cites.
Atom: Low-bit quantization for efficient and accurate llm serving
Yilong Zhao, Chien-Yu Lin, Kan Zhu, Zihao Ye, Lequn Chen, Size Zheng, Luis Ceze, Arvind Krishnamurthy, Tianqi Chen, and Baris Kasikci · 2023
Later among the works it cites.
Fluctuation-based adaptive structured pruning for large language models
Yongqi An, Xu Zhao, Tao Yu, Ming Tang, and Jinqiao Wang · 2024
Closest in time.
Scaling sparse fine-tuning to large language models
Alan Ansell, Ivan Vulić, Hannah Sterz, Anna Korhonen, and Edoardo M Ponti · 2024
Closest in time.
Quarot: Outlier-free 4-bit inference in rotated llms
Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L Croci, Bo Li, Martin Jaggi, Dan Alistarh, Torsten Hoefler, and James Hensman · 2024
Closest in time.
A survey on mixture of experts
Weilin Cai, Juyong Jiang, Fan Wang, Jing Tang, Sunghun Kim, and Jiayi Huang · 2024
Closest in time.
Shaoxiang Chen, Zequn Jie, and Lin Ma · 2024
Closest in time.
Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models, 2024
Damai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y. Wu, Zhenda Xie, Y. K. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, and Wenfeng Liang · 2024
Closest in time.
Pruner-zero: Evolving symbolic pruning metric from scratch for large language models
Peijie Dong, Lujun Li, Zhenheng Tang, Xiang Liu, Xinglin Pan, Qiang Wang, and Xiaowen Chu · 2024
Closest in time.
Mogu: A framework for enhancing safety of open-sourced llms while preserving their usability, 2024
Yanrui Du, Sendong Zhao, Danyang Zhao, Ming Ma, Yuhan Chen, Liangyu Huo, Qing Yang, Dongliang Xu, and Bing Qin · 2024
Closest in time.
Mixture of cluster-conditional lora experts for vision-language instruction tuning, 2024
Yunhao Gou, Zhili Liu, Kai Chen, Lanqing Hong, Hang Xu, Aoxue Li, Dit-Yan Yeung, James T. Kwok, and Yu Zhang · 2024
Closest in time.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2024
Closest in time.
Yefei He, Feng Chen, Jing Liu, et al · 2024
Closest in time.
Billm: Pushing the limit of post-training quantization for llms
Wei Huang, Yangdong Liu, Haotong Qin, Ying Li, Shiming Zhang, Xianglong Liu, Michele Magno, and Xiaojuan Qi · 2024
Closest in time.
Pre-gated moe: An algorithm-system co-design for fast and scalable mixture-of-expert inference
Ranggi Hwang, Jianyu Wei, Shijie Cao, Changho Hwang, Xiaohu Tang, Ting Cao, and Mao Yang · 2024
Closest in time.
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2024
Closest in time.
{ \{ InfiniGen } \} : Efficient generative inference of large language models with dynamic { \{ KV } \} cache management
Wonbeom Lee, Jungi Lee, Junghwan Seo, and Jaewoong Sim · 2024
Closest in time.
Fast and efficient 2-bit llm inference on gpu: 2/4/16-bit in a weight matrix with asynchronous dequantization
Jinhao Li, Jiaming Xu13, Shiyao Li23, Shan Huang, Jun Liu, Yaoxiu Lian, and Guohao Dai13 · 2024
Closest in time.
Jamba: A hybrid transformer-mamba language model, 2024
Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov, Shai Shalev-Shwartz, Omri Abend, Raz Alon, Tomer Asida, Amir Bergman, Roman Glozman, Michael Gokhman, Avashalom Manevich, Nir Ratner, Noam Rozen, Erez Shwartz, Mor Zusman, and Yoav Shoham · 2024
Closest in time.
Moe-llava: Mixture of experts for large vision-language models, 2024
Bin Lin, Zhenyu Tang, Yang Ye, Jinfa Huang, Junwu Zhang, Yatian Pang, Peng Jin, Munan Ning, Jiebo Luo, and Li Yuan · 2024
Closest in time.
Relu strikes back: Exploiting activation sparsity in large language models
Seyed Iman Mirzadeh, Keivan Alizadeh-Vahid, Sachin Mehta, Carlo C del Mundo, Oncel Tuzel, Golnoosh Samei, Mohammad Rastegari, and Mehrdad Farajtabar · 2024
Closest in time.
Lisa: Layerwise importance sampling for memory-efficient large language model fine-tuning
Rui Pan, Xiang Liu, Shizhe Diao, Renjie Pi, Jipeng Zhang, Chi Han, and Tong Zhang · 2024
Closest in time.
From sparse to soft mixtures of experts
Joan Puigcerver, Carlos Riquelme Ruiz, Basil Mustafa, and Neil Houlsby · 2024
Closest in time.
On the efficacy of eviction policy for key-value constrained generative language model inference
Siyu Ren and Kenny Q. Zhu · 2024
Closest in time.
Language-specific pruning for efficient reduction of large language models
Maksym Shamrai · 2024
Closest in time.
One-shot sensitivity-aware mixed sparsity pruning for large language models
Hang Shao, Bei Liu, and Yanmin Qian · 2024
Closest in time.
Efficient post-training quantization with fp8 formats
Y. Shen, X. Zhang, and J. Lin · 2024
Closest in time.
Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling
Shaohuai Shi, Xinglin Pan, Qiang Wang, Chengjian Liu, Xiaozhe Ren, Zhongzhe Hu, Yu Yang, Bo Li, and Xiaowen Chu · 2024
Closest in time.
Quest: Query-aware sparsity for efficient long-context llm inference
Jiaming Tang, Yilong Zhao, Kan Zhu, Guangxuan Xiao, Baris Kasikci, and Song Han · 2024
Closest in time.
Scaling laws with vocabulary: Larger models deserve larger vocabularies
Chaofan Tao, Qian Liu, Longxu Dou, Niklas Muennighoff, Zhongwei Wan, Ping Luo, Min Lin, and Ngai Wong · 2024
Closest in time.
Data-efficient multimodal fusion on a single gpu, 2024
Noël Vouitsis, Zhaoyan Liu, Satya Krishna Gorti, Valentin Villecroze, Jesse C. Cresswell, Guangwei Yu, Gabriel Loaiza-Ganem, and Maksims Volkovs · 2024
Closest in time.
An empirical study of mamba-based language models
Roger Waleffe, Wonmin Byeon, Duncan Riach, Brandon Norick, Vijay Korthikanti, Tri Dao, Albert Gu, et al · 2024
Closest in time.
Loma: Lossless compressed memory attention
Yumeng Wang and Zhenyang Xiao · 2024
Closest in time.
Skywork-moe: A deep dive into training techniques for mixture-of-experts language models, 2024
Tianwen Wei, Bo Zhu, Liang Zhao, Cheng Cheng, Biye Li, Weiwei Lü, Peng Cheng, Jianhao Zhang, Xiaoyu Zhang, Liang Zeng, Xiaokun Wang, Yutuan Ma, Rui Hu, Shuicheng Yan, Han Fang, and Yahui Zhou · 2024
Closest in time.
Efficient vision-language models by summarizing visual tokens into compact registers
Yuxin Wen, Qingqing Cao, Qichen Fu, et al · 2024
Closest in time.
Lamini-lm: A diverse herd of distilled models from large-scale instructions
Minghao Wu, Abdul Waheed, Chiyu Zhang, Muhammad Abdul-Mageed, and Alham Fikri Aji · 2024
Closest in time.
Haocheng Xi, Yuxiang Chen, Kang Zhao, Kai Jun Teh, Jianfei Chen, and Jun Zhu · 2024
Closest in time.
Pyramiddrop: Accelerating your large vision-language models via pyramid visual redundancy reduction
Long Xing, Qidong Huang, Xiaoyi Dong, et al · 2024
Closest in time.
Besa: Pruning large language models with blockwise parameter-efficient sparsity allocation
Peng Xu, Wenqi Shao, Mengzhao Chen, Shitao Tang, Kaipeng Zhang, Peng Gao, Fengwei An, Yu Qiao, and Ping Luo · 2024
Closest in time.
Exploiting inter-layer expert affinity for accelerating mixture-of-experts model inference
Jinghan Yao, Quentin Anthony, Aamir Shafi, Hari Subramoni, and Dhabaleswar K. DK Panda · 2024
Closest in time.
Structured pruning for large language models using coupled components elimination and minor fine-tuning
Honghe Zhang, XiaolongShi XiaolongShi, Jingwei Sun, and Guangzhong Sun · 2024
Closest in time.
Loraprune: Structured pruning meets low-rank parameter-efficient fine-tuning
Mingyang Zhang, Hao Chen, Chunhua Shen, Zhen Yang, Linlin Ou, Xinyi Yu, and Bohan Zhuang · 2024
Closest in time.
A survey on efficient inference for large language models
Zixuan Zhou, Xuefei Ning, Ke Hong, Tianyu Fu, Jiaming Xu, Shiyao Li, Yuming Lou, Luning Wang, Zhihang Yuan, Xiuhong Li, Shengen Yan, Guohao Dai, Xiao-Ping Zhang, Yuhan Dong, and Yu Wang · 2024
Closest in time.
Qrazor: Reliable and effortless 4-bit llm quantization by significant data razoring
Dongyoung Lee, Seungkyu Choi, and Ik Joon Chang · 2025
Closest in time.
Exploiting sparsity for long context inference: Million token contexts on commodity gpus
Ryan Synk, Monte Hoover, John Kirchenbauer, Neel Jain, Alex Stein, Manli Shu, Josue Melendez Sanchez, Ramani Duraiswami, and Tom Goldstein · 2025
Closest in time.
Easyspec: Layer-parallel speculative decoding for efficient multi-gpu utilization, 2025
Yize Wu, Ke Gao, and Yanjun Wu · 2025
Closest in time.
A hessian-informed hyperparameter optimization for differential learning rate
Shiyun Xu, Zhiqi Bu, Yiliang Zhang, and Ian Barnett · 2025
Closest in time.