Fetching the paper…
Reading the bibliography…
The increasing size of deep neural networks (DNNs) necessitates effective model compression to reduce their computational and memory footprints.
Optimal brain damage
Yann LeCun, John Denker, and Sara Solla · 1989
Earlier work this paper cites.
Optimal brain surgeon and general network pruning
Babak Hassibi, David G Stork, and Gregory J Wolff · 1993
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Dynamic network surgery for efficient dnns
Yiwen Guo, Anbang Yao, and Yurong Chen · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Learning structured sparsity in deep neural networks
Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Earlier work this paper cites.
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2019
Earlier work this paper cites.
Accelerator-aware pruning for convolutional neural networks
Hyeong-Ju Kang · 2019
Earlier work this paper cites.
Squantizer: Simultaneous learning for both sparse and low-precision neural networks, 2019
Mi Sun Park, Xiaofan Xu, and Cormac Brick · 2019
Earlier work this paper cites.
BFloat16: The secret to high performance on Cloud TPUs, August 2019
Shibo Wang and Pankaj Kanwar · 2019
Earlier work this paper cites.
Balanced sparsity for efficient DNN inference on GPU
Zhuliang Yao, Shijie Cao, Wencong Xiao, Chen Zhang, and Lanshun Nie · 2019
Earlier work this paper cites.
Q8BERT: quantized 8bit BERT
Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat · 2019
Earlier work this paper cites.
Language Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Pushing the limits of narrow precision inferencing at cloud scale with microsoft floating point
Bita Darvish Rouhani, Daniel Lo, Ritchie Zhao, Ming Liu, Jeremy Fowers, Kalin Ovtcharov, Anna Vinogradsky, Sarah Massengill, Lita Yang, Ray Bittner, Alessandro Forin, Haishan Zhu, Taesik Na, Prerak Patel, Shuai Che, Lok Chand Koppaka, XIA SONG, Subhojit Som, Kaustav Das, Saurabh T, Steve Reinhardt, Sitaram Lanka, Eric Chung, and Doug Burger · 2020
Earlier work this paper cites.
Rigging the lottery: Making all tickets winners
Utku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro, and Erich Elsen · 2020
Earlier work this paper cites.
Train large, then compress: Rethinking model size for efficient training and inference of transformers
Zhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin, Kurt Keutzer, Dan Klein, and Joey Gonzalez · 2020
Earlier work this paper cites.
Movement pruning: Adaptive sparsity by fine-tuning
Victor Sanh, Thomas Wolf, and Alexander Rush · 2020
Earlier work this paper cites.
Efficient Processing of Deep Neural Networks
Vivienne Sze, Yu-Hsin Chen, Tien-Ju Yang, and Joel S. Emer · 2020
Earlier work this paper cites.
Apq: Joint search for network architecture, pruning and quantization policy
Tianzhe Wang, Kuan Wang, Han Cai, Ji Lin, Zhijian Liu, Hanrui Wang, Yujun Lin, and Song Han · 2020
Earlier work this paper cites.
Joint pruning & quantization for extremely sparse neural networks
Po-Hsiang Yu, Sih-Sian Wu, Jan P Klopp, Liang-Gee Chen, and Shao-Yi Chien · 2020
Earlier work this paper cites.
Vs-quant: Per-vector scaled quantization for accurate low-precision neural network inference
Steve Dai, Rangha Venkatesan, Mark Ren, Brian Zimmer, William Dally, and Brucek Khailany · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Earlier work this paper cites.
Ps and qs: Quantization-aware pruning for efficient low latency neural network inference
Benjamin Hawks, Javier Duarte, Nicholas J Fraser, Alessandro Pappalardo, Nhan Tran, and Yaman Umuroglu · 2021
Earlier work this paper cites.
Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste · 2021
Earlier work this paper cites.
Opq: Compressing deep neural networks with one-shot pruning-quantization
Peng Hu, Xi Peng, Hongyuan Zhu, Mohamed M Sabry Aly, and Jie Lin · 2021
Cited alongside, same era.
Accelerated sparse neural training: A provable and efficient method to find N: M transposable masks
Itay Hubara, Brian Chmiel, Moshe Island, Ron Banner, Joseph Naor, and Daniel Soudry · 2021
Cited alongside, same era.
Neural network compression framework for fast model inference
Alexander Kozlov, Ivan Lazarevich, Vasily Shamporov, Nikolay Lyalyushkin, and Yury Gorbachev · 2021
Cited alongside, same era.
Nvidia ampere architecture whitepaper
Nvidia · 2021
Cited alongside, same era.
Channel permutations for N: M sparsity
Jeff Pool and Chong Yu · 2021
Cited alongside, same era.
Dominosearch: Find layer-wise fine-grained N: M sparse schemes from dense neural networks
Wei Sun, Aojun Zhou, Sander Stuijk, Rob G. J. Wijnhoven, Andrew Nelson, Hongsheng Li, and Henk Corporaal · 2021
Scaling laws for sparsely-connected foundation models
Elias Frantar, Carlos Riquelme, Neil Houlsby, Dan Alistarh, and Utku Evci · 2023
Later among the works it cites.
Enabling next-generation ai workloads: Announcing tpu v5p and ai hypercomputer
Google · 2023
Later among the works it cites.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Later among the works it cites.
Pruning vs quantization: Which is better?
Andrey Kuzmin, Markus Nagel, Mart van Baalen, Arash Behboodi, and Tijmen Blankevoort · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Training deep neural networks with joint quantization and pruning of weights and activations
Xinyu Zhang, Ian Colbert, Ken Kreutz-Delgado, and Srinjoy Das · 2021
Cited alongside, same era.
Learning N: M fine-grained structured sparse neural networks from scratch
Aojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu, Zhijie Zhang, Kun Yuan, Wenxiu Sun, and Hongsheng Li · 2021
Cited alongside, same era.
Llm.int8(): 8-bit matrix multiplication for transformers at scale
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer · 2022
Cited alongside, same era.
Optimal brain compression: A framework for accurate post-training quantization and pruning
Elias Frantar and Dan Alistarh · 2022
Cited alongside, same era.
GPTQ: accurate post-training quantization for generative pre-trained transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh · 2022
Cited alongside, same era.
Accuracy boosters: Epoch-driven mixed-mantissa block floating-point for dnn training
Simla Burcu Harma, Ayan Chakraborty, Babak Falsafi, Martin Jaggi, and Yunho Oh · 2022
Cited alongside, same era.
Mike Lasby, Anna Golubeva, Utku Evci, Mihai Nica, and Yani Ioannou · 2023
Later among the works it cites.
Deja vu: Contextual sparsity for efficient llms at inference time
Zichang Liu, Jue Wang, Tri Dao, Tianyi Zhou, Binhang Yuan, Zhao Song, Anshumali Shrivastava, Ce Zhang, Yuandong Tian, Christopher Re, et al · 2023
Later among the works it cites.
STEP: learning N: M structured sparsity masks from scratch with precondition
Yucheng Lu, Shivani Agrawal, Suvinay Subramanian, Oleg Rybakov, Christopher De Sa, and Amir Yazdanbakhsh · 2023
Later among the works it cites.
OCP 8-bit Floating Point Specification (OFP8)
Paulius Micikevicius, Stuart Oberman, Pradeep Dubey, Marius Cornea, Andres Rodriguez, Ian Bratt, Richard Grisenthwaite, Norm Jouppi, Chiachen Chou, Amber Huffman, Michael Schulte, Ralph Wittig, Dharmesh Jani, and Summer Deng · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
With shared microexponents, A little shifting goes a long way
Bita Darvish Rouhani, Ritchie Zhao, Venmugil Elango, Rasoul Shafipour, Mathew Hall, Maral Mesmakhosroshahi, Ankit More, Levi Melnick, Maximilian Golub, Girish Varatkar, Lai Shao, Gaurav Kolhe, Dimitry Melts, Jasmine Klar, Renee L’Heureux, Matt Perry, Doug Burger, Eric S. Chung, Zhaoxia (Summer) Deng, Sam Naghshineh, Jongsoo Park, and Maxim Naumov · 2023
Later among the works it cites.
A simple and effective pruning approach for large language models
Mingjie Sun, Zhuang Liu, Anna Bair, and J. Zico Kolter · 2023
Later among the works it cites.
Prune and tune: Improving efficient pruning techniques for massive language models
Aaquib Syed, Phillip Huang Guo, and Vijaykaarti Sundarapandiyan · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton-Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurélien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom · 2023
Later among the works it cites.
FP8 versus INT8 for efficient deep learning inference
Mart van Baalen, Andrey Kuzmin, Suparna S. Nair, Yuwei Ren, Eric Mahurin, Chirag Patel, Sundar Subramanian, Sanghyuk Lee, Markus Nagel, Joseph Soriaga, and Tijmen Blankevoort · 2023
Later among the works it cites.
Understanding int4 quantization for transformer models: Latency speedup, composability, and failure cases, 2023
Xiaoxia Wu, Cheng Li, Reza Yazdani Aminabadi, Zhewei Yao, and Yuxiong He · 2023
Later among the works it cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Guangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu, Julien Demouth, and Song Han · 2023
Later among the works it cites.
Boost vision transformer with gpu-friendly sparsity and quantization
Chong Yu, Tao Chen, Zhongxue Gan, and Jiayuan Fan · 2023
Later among the works it cites.
Towards optimal compression: Joint pruning and quantization
Ben Zandonati, Glenn Bucagu, Adrian Alan Pol, Maurizio Pierini, Olya Sirkin, and Tal Kopetz · 2023
Later among the works it cites.
Progressive gradient flow for robust N: M sparsity training in transformers
Abhimanyu Rajeshkumar Bambhaniya, Amir Yazdanbakhsh, Suvinay Subramanian, Sheng-Chun Kao, Shivani Agrawal, Utku Evci, and Tushar Krishna · 2024
Closest in time.
A framework for few-shot language model evaluation, 07 2024
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou · 2024
Closest in time.
OWQ: outlier-aware weight quantization for efficient fine-tuning and inference of large language models
Changhun Lee, Jungyu Jin, Taesu Kim, Hyungjun Kim, and Eunhyeok Park · 2024
Closest in time.
SPP: sparsity-preserved parameter-efficient fine-tuning for large language models
Xudong Lu, Aojun Zhou, Yuhui Xu, Renrui Zhang, Peng Gao, and Hongsheng Li · 2024
Closest in time.
The era of 1-bit llms: All large language models are in 1.58 bits, 2024
Shuming Ma, Hongyu Wang, Lingxiao Ma, Lei Wang, Wenhui Wang, Shaohan Huang, Li Dong, Ruiping Wang, Jilong Xue, and Furu Wei · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex Castro-Ros, Ambrose Slone, Amélie Héliou, Andrea Tacchetti, Anna Bulanova, Antonia Paterson, Beth Tsai, Bobak Shahriari, Charline Le Lan, Christopher A. Choquette-Choo, Clément Crepy, Daniel Cer, Daphne Ippolito, David Reid, Elena Buchatskaya, Eric Ni, Eric Noland, Geng Yan, George Tucker, George-Christian Muraru, Grigory Rozhdestvenskiy, Henryk Michalewski, Ian Tenney, Ivan Grishchenko, Jacob Austin, James Keeling, Jane Labanowski, Jean-Baptiste Lespiau, Jeff Stanway, Jenny Brennan, Jeremy Chen, Johan Ferret, Justin Chiu, and et al · 2024
Closest in time.
MicroXcaling: A Library for Microservices Autoscaling
Microsoft · 2024
Closest in time.
Nvidia tensor cores: Unprecedented acceleration for generative ai
Nvidia · 2024
Closest in time.