Fetching the paper…
Reading the bibliography…
Deploying Large Language Models (LLMs) on edge or mobile devices offers significant benefits, such as enhanced data privacy and real-time processing capabilities.
C. E. Shannon, “A mathematical theory of communication,” The Bell system technical journal
1948
Earlier work this paper cites.
D. A. Huffman, “A method for the construction of minimum-redundancy codes,” Proceedings of the IRE
1952
Earlier work this paper cites.
J. Ziv and A. Lempel, “A universal algorithm for sequential data compression,” IEEE Transactions on information theory
1977
Earlier work this paper cites.
P. Deutsch, “Deflate compressed data format specification version 1.3,” tech. rep., 1996
1996
Earlier work this paper cites.
P. Deutsch, “GZIP file format specification version 4.3,” RFC
1996
Earlier work this paper cites.
2013
Earlier work this paper cites.
2015
Earlier work this paper cites.
Y. Collet, “Zstd github repository from facebook.,” 2016
2016
Earlier work this paper cites.
P. Jain, A. Jain, A. Nrusimha, A. Gholami, P. Abbeel, J. Gonzalez, K. Keutzer, and I. Stoica, “Checkmate: Breaking the memory wall with optimal tensor rematerialization,” Proceedings of Machine Learning and Systems
2020
Earlier work this paper cites.
J. Ren, S. Rajbhandari, R. Y. Aminabadi, O. Ruwase, S. Yang, M. Zhang, D. Li, and Y. He, “ { \{ ZeRO-Offload } \} : Democratizing { \{ Billion-Scale } \} model training,” in 2021 USENIX Annual Technical Conference (USENIX ATC 21)
2021
Earlier work this paper cites.
T. Dao, D. Fu, S. Ermon, A. Rudra, and C. Ré, “Flashattention: Fast and memory-efficient exact attention with io-awareness,” Advances in Neural Information Processing Systems
2022
Earlier work this paper cites.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Z. Yao, R. Yazdani Aminabadi, M. Zhang, X. Wu, C. Li, and Y. He, “Zeroquant: Efficient and affordable post-training quantization for large-scale transformers,” Advances in Neural Information Processing Systems
2022
Cited alongside, same era.
2023
Later among the works it cites.
G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han, “Smoothquant: Accurate and efficient post-training quantization for large language models,” in International Conference on Machine Learning
2023
Later among the works it cites.
2023
Later among the works it cites.
V. A. Korthikanti, J. Casper, S. Lym, L. McAfee, M. Andersch, M. Shoeybi, and B. Catanzaro, “Reducing activation recomputation in large transformer models,” Proceedings of Machine Learning and Systems
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
Y. Mao et al
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Y. Sheng, L. Zheng, B. Yuan, Z. Li, M. Ryabinin, B. Chen, P. Liang, C. Ré, I. Stoica, and C. Zhang, “Flexgen: High-throughput generative inference of large language models with a single gpu,” in International Conference on Machine Learning
2023
Cited alongside, same era.
2023
Cited alongside, same era.
L. Zhang, X. Liu, Z. Li, X. Pan, P. Dong, R. Fan, R. Guo, X. Wang, Q. Luo, S. Shi, et al
2023
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Mao, J. Li, Y. Cui, and J. C. Xue, “Faster and stronger lossless compression with optimized autoregressive framework,” in 2023 60th ACM/IEEE Design Automation Conference (DAC)
2023
Later among the works it cites.
J. Rae, “Compression for agi,” tech. rep., 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Yu, S. Buchanan, D. Pai, T. Chu, Z. Wu, S. Tong, B. Haeffele, and Y. Ma, “White-box transformers via sparse rate reduction,” Advances in Neural Information Processing Systems
2024
Closest in time.