Fetching the paper…
Reading the bibliography…
With the growth of model sizes and scale of their deployment, their sheer size burdens the infrastructure requiring more network and more storage to accommodate these.
D. A. Huffman, “A method for the construction of minimum-redundancy codes,” Proceedings of the IRE
1952
Earlier work this paper cites.
J. J. Rissanen, “Generalized kraft inequality and arithmetic coding,” IBM Journal of Research and Development
1976
Earlier work this paper cites.
J. Ziv and A. Lempel, “A universal algorithm for sequential data compression,” IEEE Transactions on Information Theory
1977
Earlier work this paper cites.
R. Gray, “Vector quantization,” IEEE Assp Magazine
1984
Earlier work this paper cites.
S. J. Hanson and L. Y. Pratt, “Comparing biases for minimal network construction with back-propagation,” in Neural Information Processing Systems
1988
Earlier work this paper cites.
Y. LeCun, J. S. Denker, and S. A. Solla, “Optimal brain damage,” in Neural Information Processing Systems
1989
Earlier work this paper cites.
P. Deutsch and J.-L. Gailly, “Zlib compressed data format specification version 3.3,” tech. rep., 1996
1996
Earlier work this paper cites.
C.-Y. Lin, “ROUGE: A package for automatic evaluation of summaries,” in Text Summarization Branches Out
2004
Earlier work this paper cites.
B. Pang and L. Lee, “Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales,” in Proceedings of the ACL
2005
Earlier work this paper cites.
A. Fader, L. Zettlemoyer, and O. Etzioni, “Open Question Answering Over Curated and Extracted Knowledge Bases,” in KDD
2014
Earlier work this paper cites.
Software available from tensorflow.org
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng, “TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015 · 2015
Earlier work this paper cites.
F. Chollet et al
2015
Earlier work this paper cites.
S. Han, H. Mao, and W. J. Dally, “Deep compression: Compressing deep neural network with pruning, trained quantization and huffman coding,” arXiv: Computer Vision and Pattern Recognition
2015
Earlier work this paper cites.
Squash, “Squash compression benchmark,” 2016
2016
Earlier work this paper cites.
R. Nallapati, B. Zhou, C. Gulcehre, B. Xiang, et al
2016
Earlier work this paper cites.
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang, “SQuAD: 100,000+ questions for machine comprehension of text,” in Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
Y. Collet and M. Kucherawy, “Zstandard compression and the application/zstd media type,” tech. rep., 2018
2018
Earlier work this paper cites.
J. Yu Koh, “Model zoo (hub),” 2018
2018
Earlier work this paper cites.
Google, “Tensorflow hub,” 2018
2018
Earlier work this paper cites.
M. Junczys-Dowmunt, R. Grundkiewicz, T. Dwojak, H. T. Hoang, K. Heafield, T. Neckermann, F. Seide, U. Germann, A. F. Aji, N. Bogoychev, A. F. T. Martins, and A. Birch, “Marian: Fast neural machine translation in c++,” in Annual Meeting of the Association for Computational Linguistics
2018
Earlier work this paper cites.
S. Narayan, S. B. Cohen, and M. Lapata, “Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing
2018
Earlier work this paper cites.
M. Post, “A call for clarity in reporting BLEU scores,” in Proceedings of the Third Conference on Machine Translation: Research Papers
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in North American Chapter of the Association for Computational Linguistics
2019
Earlier work this paper cites.
Pytorch, “Pytorch hub,” 2019
2019
Earlier work this paper cites.
S. Wang and P. Kanwar, “Bfloat16: The secret to high performance on cloud tpus,” Google Cloud Blog
2019
Earlier work this paper cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “Pytorch: An imperative style, high-performance deep learning library,” in Advances in Neural Information Processing Systems 32
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
M. Haroush, I. Hubara, E. Hoffer, and D. Soudry, “The knowledge within: Methods for data-free model compression,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2022
Later among the works it cites.
M. Wortsman, G. Ilharco, S. Y. Gadre, R. Roelofs, R. Gontijo-Lopes, A. S. Morcos, H. Namkoong, A. Farhadi, Y. Carmon, S. Kornblith, and L. Schmidt, “Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time,” 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
A. Radford, J. Wu, D. Amodei, D. Amodei, J. Clark, M. Brundage, and I. Sutskever, “Better language models and their implications,” OpenAI blog
2019
Cited alongside, same era.
S. Schneider, A. Baevski, R. Collobert, and M. Auli, “wav2vec: Unsupervised Pre-Training for Speech Recognition,” in Proc. Interspeech 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
J. Pfeiffer, A. Rücklé, C. Poth, A. Kamath, I. Vulić, S. Ruder, K. Cho, and I. Gurevych, “AdapterHub: A framework for adapting transformers,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations
2020
Cited alongside, same era.
2020
Cited alongside, same era.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” The Journal of Machine Learning Research
2020
Cited alongside, same era.
2020
Cited alongside, same era.
T. Choudhary, V. Mishra, A. Goswami, and J. Sarangapani, “A comprehensive survey on model compression and acceleration,” Artificial Intelligence Review
2020
Cited alongside, same era.
2023
Later among the works it cites.
S. Don-Yehiya, E. Venezian, C. Raffel, N. Slonim, and L. Choshen, “ColD fusion: Collaborative descent for distributed multitask finetuning,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
V. Lialin, S. Muckatira, N. Shivagunde, and A. Rumshisky, “Relora: High-rank training through low-rank updates,” in Workshop on Advancing Neural Network Training: Computational Efficiency, Scalability, and Resource Optimization (WANT@ NeurIPS 2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Z. Liu, A. Qiao, W. Neiswanger, H. Wang, B. Tan, T. Tao, J. Li, Y. Wang, S. Sun, O. Pangarkar, R. Fan, Y. Gu, V. Miller, Y. Zhuang, G. He, H. Li, F. Koto, L. Tang, N. Ranjan, Z. Shen, X. Ren, R. Iriondo, C. Mu, Z. Hu, M. Schulze, P. Nakov, T. Baldwin, and E. P. Xing, “Llm360: Towards fully transparent open-source llms,” arXiv
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
A. Gueta, E. Venezian, C. Raffel, N. Slonim, Y. Katz, and L. Choshen, “Knowledge is a region in weight space for fine-tuned language models,” in Findings of the Association for Computational Linguistics: EMNLP 2023
2023
Later among the works it cites.
P. Yadav, D. Tam, L. Choshen, C. Raffel, and M. Bansal, “Ties-merging: Resolving interference when merging models,” in Thirty-seventh Conference on Neural Information Processing Systems
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
P. Wu, “Pytorch 2.0: The journey to bringing compiler technologies to the core of pytorch (keynote),” in Proceedings of the 21st ACM/IEEE International Symposium on Code Generation and Optimization
2023
Later among the works it cites.
S. Tyagi and M. Swany, “Gravac: Adaptive compression for communication-efficient distributed dl training,” in 2023 IEEE 16th International Conference on Cloud Computing (CLOUD)
2023
Later among the works it cites.
2023
Later among the works it cites.
X. Geng and H. Liu, “Openllama: An open reproduction of llama,” May 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
H. Ivison, Y. Wang, V. Pyatkin, N. Lambert, M. Peters, P. Dasigi, J. Jang, D. Wadden, N. A. Smith, I. Beltagy, and H. Hajishirzi, “Camels in a changing climate: Enhancing lm adaptation with tulu 2,” 2023
2023
Later among the works it cites.
Y. Collet, “Lz4 - extremely fast compression,” 2024
2024
Closest in time.
M. Adler and J.-L. Gailly, “Zlib,” 2024
2024
Closest in time.
Y. Collet, “Zstandard,” 2024
2024
Closest in time.