Fetching the paper…
Reading the bibliography…
The high computational and memory requirements of large language model (LLM) inference make it feasible only with multiple high-end accelerators.
I/o complexity: The red-blue pebble game
Jia-Wei, H. and Kung, H.-T · 1981
Earlier work this paper cites.
Pagecachemangagement
Morton, A · 2008
Earlier work this paper cites.
Communication-avoiding algorithms for linear algebra and beyond
Demmel, J · 2013
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context
Paperno, D., Kruszewski, G., Lazaridou, A., Pham, N.-Q., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R · 2016
Earlier work this paper cites.
Narayan, S., Cohen, S. B., and Lapata, M · 2018
Earlier work this paper cites.
Superneurons: Dynamic gpu memory management for training deep neural networks
Wang, L., Ye, J., Zhao, Y., Wu, W., Li, A., Song, S. L., Xu, Z., and Kraska, T · 2018
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., et al · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Swapadvisor: Pushing deep learning beyond the gpu memory limit via smart swapping
Huang, C.-C., Jin, G., and Li, J · 2020
Earlier work this paper cites.
Q-bert: Hessian based ultra low precision quantization of bert
Shen, S., Dong, Z., Ye, J., Ma, L., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K · 2020
Earlier work this paper cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Earlier work this paper cites.
Spreadsheetcoder: Formula prediction from semi-structured context
Chen, X., Maniatis, P., Singh, R., Sutton, C., Dai, H., Lin, M., and Zhou, D · 2021
Earlier work this paper cites.
Turbotransformers: an efficient gpu serving system for transformer models
Fang, J., Yu, Y., Zhao, C., and Zhou, J · 2021
Earlier work this paper cites.
Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks
Hoefler, T., Alistarh, D., Ben-Nun, T., Dryden, N., and Peste, A · 2021
Cited alongside, same era.
Efficient large-scale language model training on gpu clusters using megatron-lm
Narayanan, D., Shoeybi, M., Casper, J., LeGresley, P., Patwary, M., Korthikanti, V., Vainbrand, D., Kashinkunti, P., Bernauer, J., Catanzaro, B., et al · 2021
Cited alongside, same era.
Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning
Rajbhandari, S., Ruwase, O., Rasley, J., Smith, S., and He, Y · 2021
Cited alongside, same era.
Zero-offload: Democratizing billion-scale model training
Ren, J., Rajbhandari, S., Aminabadi, R. Y., Ruwase, O., Yang, S., Zhang, M., Li, D., and He, Y · 2021
Cited alongside, same era.
Lightseq: A high performance inference library for transformers
Wang, X., Xiong, Y., Wei, Y., Wang, M., and Li, L · 2021
Cited alongside, same era.
Li, Y., Phanishayee, A., Murray, D., Tarnawski, J., and Kim, N. S · 2022
Later among the works it cites.
Holistic evaluation of language models
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., et al · 2022
Later among the works it cites.
Can foundation models wrangle your data?
Narayan, A., Chami, I., Orr, L., and Ré, C · 2022
Later among the works it cites.
Fastertransformer
NVIDIA · 2022
Later among the works it cites.
nuqmm: Quantized matmul for efficient inference of large-scale generative language models
Park, G., Park, B., Kwon, S. J., Kim, B., Lee, Y., and Lee, D · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deepspeed-inference: Enabling efficient inference of transformer models at unprecedented scale
Aminabadi, R. Y., Rajbhandari, S., Awan, A. A., Li, C., Li, D., Zheng, E., Ruwase, O., Smith, S., Zhang, M., Rasley, J., et al · 2022
Cited alongside, same era.
Petals: Collaborative inference and fine-tuning of large models
Borzunov, A., Baranchuk, D., Dettmers, T., Ryabinin, M., Belkada, Y., Chumachenko, A., Samygin, P., and Raffel, C · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Cited alongside, same era.
The case for 4-bit precision: k-bit inference scaling laws
Dettmers, T. and Zettlemoyer, L · 2022
Cited alongside, same era.
Llm.int8(): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2022
Cited alongside, same era.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2022
Cited alongside, same era.
Later among the works it cites.
Efficiently scaling transformer inference
Pope, R., Douglas, S., Chowdhery, A., Devlin, J., Bradbury, J., Levskaya, A., Heek, J., Xiao, K., Agrawal, S., and Dean, J · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
Scao, T. L., Fan, A., Akiki, C., Pavlick, E., Ilić, S., Hesslow, D., Castagné, R., Luccioni, A. S., Yvon, F., Gallé, M., et al · 2022
Later among the works it cites.
Olla: Optimizing the lifetime and location of arrays to reduce the memory usage of neural networks
Steiner, B., Elhoushi, M., Kahn, J., and Hegarty, J · 2022
Later among the works it cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Demouth, J., and Han, S · 2022
Later among the works it cites.
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Yao, Z., Aminabadi, R. Y., Zhang, M., Wu, X., Li, C., and He, Y · 2022
Later among the works it cites.
Orca: A distributed serving system for { \{ Transformer-Based } \} generative models
Yu, G.-I., Jeong, J. S., Kim, G.-W., Kim, S., and Chun, B.-G · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Later among the works it cites.
Alpa: Automating inter-and intra-operator parallelism for distributed deep learning
Zheng, L., Li, Z., Zhang, H., Zhuang, Y., Chen, Z., Huang, Y., Wang, Y., Xu, Y., Zhuo, D., Gonzalez, J. E., et al · 2022
Later among the works it cites.
Massive language models can be accurately pruned in one-shot
Frantar, E. and Alistarh, D · 2023
Closest in time.
Swarm parallelism: Training large models can be surprisingly communication-efficient
Ryabinin, M., Dettmers, T., Diskin, M., and Borzunov, A · 2023
Closest in time.