Fetching the paper…
Reading the bibliography…
Large language model (LLM) pruning with fixed N:M structured sparsity significantly limits the expressivity of the sparse model, yielding sub-optimal performance.
2015
Earlier work this paper cites.
H. Sharma, J. Park, E. Amaro, B. Thwaites, P. Kotha, A. Gupta, J. K. Kim, A. Mishra, and H. Esmaeilzadeh, “Dnnweaver: From high-level deep network models to fpga acceleration,” in
2016
Earlier work this paper cites.
A. Parashar, M. Rhu, A. Mukkara, A. Puglielli, R. Venkatesan, B. Khailany, J. Emer, S. W. Keckler, and W. J. Dally, “Scnn: An accelerator for compressed-sparse convolutional neural networks,”
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
F. Farshchi, Q. Huang, and H. Yun, “Integrating nvidia deep learning accelerator (nvdla) with risc-v soc on firesim,” in
2019
Earlier work this paper cites.
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach, “Towards vqa models that can read,” in
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Y. Bisk, R. Zellers, J. Gao, Y. Choi
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Q. Liu, B. Gao, P. Yao, D. Wu, J. Chen, Y. Pang, W. Zhang, Y. Liao, C.-X. Xue, W.-H. Chen
2020
Earlier work this paper cites.
Z. Chen, X. Chen, and J. Gu, “15.3 a 65nm 3t dynamic analog ram-based computing-in-memory macro and cnn accelerator with retention enhancement, adaptive analog sparsity and 44tops/w system energy efficiency,” in
2021
Earlier work this paper cites.
J.-W. Jang
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
S. Sarangi and B. Baas, “Deepscaletool: A tool for the accurate estimation of technology scaling in the deep-submicron era,” in
2021
Cited alongside, same era.
S. Yu, H. Jiang, S. Huang, X. Peng, and A. Lu, “Compute-in-memory chips for deep learning: Recent trends and prospects,”
2021
Cited alongside, same era.
H. Fujiwara
2022
Cited alongside, same era.
Z.-G. Liu, P. N. Whatmough, Y. Zhu, and M. Mattina, “S2ta: Exploiting structured sparsity for energy-efficient mobile cnn acceleration,” in
2022
Cited alongside, same era.
F. Tu, Y. Wang, L. Liang, Y. Ding, L. Liu, S. Wei, S. Yin, and Y. Xie, “Sdp: Co-designing algorithm, dataflow, and architecture for in-sram sparse nn acceleration,”
2022
Cited alongside, same era.
B. Zhong, M. Wang, C. Zhang, Y. Mai, X. Li, and Z. Yu, “A digital sram computing-in-memory design utilizing activation unstructured sparsity for high-efficient dnn inference,” in
2023
Later among the works it cites.
A. Agrawal, A. Agarwal, N. Kedia, J. Mohan, S. Kundu, N. Kwatra, R. Ramjee, and A. Tumanov, “Metron: Holistic performance evaluation framework for llm inference systems,”
2024
Later among the works it cites.
2024
Later among the works it cites.
J. Lin, J. Tang, H. Tang, S. Yang, W.-M. Chen, W.-C. Wang, G. Xiao, X. Dang, C. Gan, and S. Han, “Awq: Activation-aware weight quantization for on-device llm compression and acceleration,”
2024
Later among the works it cites.
A. Meta, “Introducing meta llama 3: The most capable openly available llm to date,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
E. Frantar and D. Alistarh, “SparseGPT: Massive language models can be accurately pruned in one-shot,” 2023
2023
Cited alongside, same era.
A. Gu and T. Dao, “Mamba: Linear-time sequence modeling with selective state spaces,”
2023
Cited alongside, same era.
G. Jeong, S. Damani, A. R. Bambhaniya, E. Qin, C. J. Hughes, S. Subramoney, H. Kim, and T. Krishna, “Vegeta: Vertically-integrated extensions for sparse/dense gemm tile acceleration on cpus,” in
2023
Cited alongside, same era.
D. Li, T. Yamasaki, A. Mani, A. T. Do, N. Chen, and B. Wang, “Laxor: A bit-accurate bnn accelerator with latch-xor logic for local computing,” in
2023
Cited alongside, same era.
H. E. Sumbul, J.-s. Seo, D. H. Morris, and E. Beigne, “A fully-digital and row-pipelined compute-in-memory neural network accelerator with soc-level benchmarking for ar/vr applications,”
2023
Cited alongside, same era.
H. Touvron
2023
Cited alongside, same era.
2024
Later among the works it cites.
A. Raha, D. A. Mathaikutty, S. K. Ghosh, and S. Kundu, “FlexNN: A dataflow-aware flexible deep learning accelerator for energy-efficient edge devices,”
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
A. Sridharan, F. Zhang, J.-S. Seo, and D. Fan, “Sp-imc: A sparsity aware in-memory-computing macro in 28nm cmos with configurable sparse representation for highly sparse dnn workloads,” in
2024
Later among the works it cites.
M. Sun, Z. Liu, A. Bair, and J. Z. Kolter, “A simple and effective pruning approach for large language models,”
2024
Later among the works it cites.
L. Yin, S. Liu, A. Jaiswal, S. Kundu, and Z. Wang, “A task-centric angle of llm pre-trained weights through sparsity,”
2024
Later among the works it cites.
L. Yin, Y. Wu, Z. Zhang, C.-Y. Hsieh, Y. Wang, Y. Jia, G. Li, A. Jaiswal, M. Pechenizkiy, Y. Liang
2024
Later among the works it cites.
Y. Zhan, W.-H. Yu, K.-F. Un, R. P. Martins, and P.-I. Mak, “A 28-nm 18.7 tops/mm2 89.4-to-234.6 tops/w 8b single-finger edram compute-in-memory macro with bit-wise sparsity aware and kernel-wise weight update/refresh,”
2024
Later among the works it cites.