Fetching the paper…
Reading the bibliography…
The field of efficient Large Language Model (LLM) inference is rapidly evolving, presenting a unique blend of opportunities and challenges.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., and Salakhutdinov, R. (2019) · 1901
Earlier work this paper cites.
Imitation learning for non-autoregressive neural machine translation
Wei, B., Wang, M., Zhou, H., Lin, J., Xie, J., and Sun, X. (2019) · 1906
Earlier work this paper cites.
Hint-based training for non-autoregressive machine translation
Li, Z., Lin, Z., He, D., Tian, F., Qin, T., Wang, L., and Liu, T.-Y. (2019) · 1909
Earlier work this paper cites.
Optimal brain damage
LeCun, Y., Denker, J., and Solla, S. (1989) · 1989
Earlier work this paper cites.
Iterative solution of nonlinear equations in several variables
Ortega, J. M. and Rheinboldt, W. C. (2000) · 2000
Earlier work this paper cites.
Addressing some limitations of transformers with feedback memory
Fan, A., Lavril, T., Grave, E., Joulin, A., and Sukhbaatar, S. (2020) · 2002
Earlier work this paper cites.
Etc: Encoding long and structured inputs in transformers
Ainslie, J., Ontanon, S., Alberti, C., Cvicek, V., Fisher, Z., Pham, P., Ravula, A., Sanghai, S., Wang, Q., and Yang, L. (2020) · 2004
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A. (2020) · 2004
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z. (2020) · 2006
Earlier work this paper cites.
Memformer: A memory-augmented transformer for sequence modeling
Wu, Q., Lan, Z., Qian, K., Gu, J., Geramifard, A., and Yu, Z. (2020) · 2010
Earlier work this paper cites.
Tensor-train decomposition
Oseledets, I. V. (2011) · 2011
Earlier work this paper cites.
Twenty years of mixture of experts
Yuksel, S. E., Wilson, J. N., and Gader, P. D. (2012) · 2012
Earlier work this paper cites.
Binaryconnect: Training deep neural networks with binary weights during propagations
Courbariaux, M., Bengio, Y., and David, J.-P. (2015) · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J. (2015) · 2015
Earlier work this paper cites.
Binarized neural networks
Hubara, I., Courbariaux, M., Soudry, D., El-Yaniv, R., and Bengio, Y. (2016) · 2016
Earlier work this paper cites.
Non-autoregressive neural machine translation
Gu, J., Bradbury, J., Xiong, C., Li, V. O., and Socher, R. (2017) · 2017
Earlier work this paper cites.
Literature survey on low rank approximation of matrices
Kishore Kumar, N. and Schneider, J. (2017) · 2017
Earlier work this paper cites.
Tvm: An automated end-to-end optimizing compiler for deep learning
Chen, T., Moreau, T., Jiang, Z., Zheng, L., Yan, E., Shen, H., Cowan, M., Wang, L., Hu, Y., Ceze, L., et al. (2018) · 2018
Earlier work this paper cites.
Pact: Parameterized clipping activation for quantized neural networks
Choi, J., Wang, Z., Venkataramani, S., Chuang, P. I.-J., Srinivasan, V., and Gopalakrishnan, K. (2018) · 2018
Earlier work this paper cites.
Deterministic non-autoregressive neural sequence modeling by iterative refinement
Lee, J., Mansimov, E., and Cho, K. (2018) · 2018
Earlier work this paper cites.
Energy-efficient neural network accelerator based on outlier-aware low-precision computation
Park, E. et al. (2018) · 2018
Earlier work this paper cites.
Blockwise parallel decoding for deep autoregressive models
Stern, M., Shazeer, N., and Uszkoreit, J. (2018) · 2018
Earlier work this paper cites.
The true processing in memory accelerator
Devaux, F. (2019a) · 2019
Earlier work this paper cites.
The true processing in memory accelerator
Devaux, F. (2019b) · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019) · 2019
Earlier work this paper cites.
Hawq: Hessian aware quantization of neural networks with mixed-precision
Dong, Z., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K. (2019) · 2019
Earlier work this paper cites.
Mask-predict: Parallel decoding of conditional masked language models
Ghazvininejad, M., Levy, O., Liu, Y., and Zettlemoyer, L. (2019) · 2019
Earlier work this paper cites.
Dynamics of deep neural networks and neural tangent hierarchy
Huang, J. and Yau, H.-T. (2019) · 2019
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., and Lillicrap, T. P. (2019) · 2019
Earlier work this paper cites.
Bert and pals: Projected attention layers for efficient adaptation in multi-task learning
Stickland, A. C. and Murray, I. (2019) · 2019
Earlier work this paper cites.
Non-autoregressive machine translation with auxiliary regularization
Wang, Y., Tian, F., He, D., Qin, T., Zhai, C., and Liu, T.-Y. (2019) · 2019
Earlier work this paper cites.
N-ode transformer: A depth-adaptive variant of the transformer using neural ordinary differential equations
Baier-Reinio, A. and Sterck, H. D. (2020) · 2020
Earlier work this paper cites.
Funnel-transformer: Filtering out sequential redundancy for efficient language processing
Dai, Z., Lai, G., Yang, Y., and Le, Q. (2020) · 2020
Earlier work this paper cites.
Depth-adaptive transformer
Elbayad, M., Gu, J., Grave, E., and Auli, M. (2020) · 2020
Earlier work this paper cites.
Shortcut learning in deep neural networks
Geirhos, R., Jacobsen, J.-H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., and Wichmann, F. A. (2020) · 2020
Earlier work this paper cites.
Fully non-autoregressive neural machine translation: Tricks of the trade
Gu, J. and Kong, X. (2020) · 2020
Earlier work this paper cites.
Jointly masked sequence-to-sequence model for non-autoregressive neural machine translation
Guo, J., Xu, L., and Chen, E. (2020) · 2020
Earlier work this paper cites.
Retrieval augmented language model pre-training
Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M. (2020) · 2020
Earlier work this paper cites.
Dynabert: Dynamic bert with adaptive width and depth
Hou, L., Huang, Z., Shang, L., Jiang, X., Chen, X., and Liu, Q. (2020) · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. (2020) · 2020
Earlier work this paper cites.
Fastbert: a self-distilling bert with adaptive inference time
Liu, W., Zhou, P., Zhao, Z., Wang, Z., Deng, H., and Ju, Q. (2020) · 2020
Earlier work this paper cites.
The right tool for the job: Matching model and instance complexities
Schwartz, R., Stanovsky, G., Swayamdipta, S., Dodge, J., and Smith, N. A. (2020) · 2020
Earlier work this paper cites.
Minimizing the bag-of-ngrams difference for non-autoregressive neural machine translation
Shao, C., Zhang, J., Feng, Y., Meng, F., and Zhou, J. (2020) · 2020
Earlier work this paper cites.
Algorithm-hardware co-design of adaptive floating-point encodings for resilient deep learning inference
Tambe, T., Yang, E.-Y., Wan, Z., Deng, Y., Reddi, V. J., Rush, A., Brooks, D., and Wei, G.-Y. (2020) · 2020
Earlier work this paper cites.
Deebert: Dynamic early exiting for accelerating bert inference
Xin, J., Tang, R., Lee, J., Yu, Y., and Lin, J. (2020) · 2020
Earlier work this paper cites.
Gobo: Quantizing attention-based nlp models for low latency and energy efficient inference
Zadeh, A. H. et al. (2020) · 2020
Earlier work this paper cites.
Bert loses patience: Fast and robust inference with early exit
Zhou, W., Xu, C., Ge, T., McAuley, J., Xu, K., and Wei, F. (2020) · 2020
Earlier work this paper cites.
Segatron: Segment-aware transformer for language modeling and understanding
Bai, H., Shi, P., Lin, J., Xie, Y., Tan, L., Xiong, K., Gao, W., and Li, M. (2021) · 2021
Earlier work this paper cites.
Improving language models by retrieving from trillions of tokens. arxiv e-prints, art
Borgeaud, S. et al. (2021) · 2021
Earlier work this paper cites.
Permuteformer: Efficient relative position encoding for long sequences
Chen, P. (2021) · 2021
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.(2021)
Fedus, W., Zoph, B., and Shazeer, N. (2021) · 2021
Earlier work this paper cites.
Knowledge distillation: A survey
Gou, J., Yu, B., Maybank, S. J., and Tao, D. (2021) · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. (2021) · 2021
Earlier work this paper cites.
25.4 A 20nm 6gb function-in-memory dram, based on HBM2 with a 1.2tflops programmable computing unit using bank-level parallelism, for machine learning applications
Kwon, Y., Lee, S. H., Lee, J., Kwon, S., Ryu, J., Son, J., O, S., Yu, H., Lee, H., Kim, S. Y., Cho, Y., Kim, J. G., Choi, J., Shin, H., Kim, J., Phuah, B., Kim, H., Song, M. J., Choi, A., Kim, D., Kim, S., Kim, E., Wang, D., Kang, S., Ro, Y., Seo, S., Song, J., Youn, J., Sohn, K., and Kim, N. S. (2021) · 2021
Earlier work this paper cites.
Alps: Adaptive quantization of deep neural networks with generalized posits
Langroudi, H. F., Karia, V., Carmichael, Z., Zyarah, A., Pandit, T., Gustafson, J. L., and Kudithipudi, D. (2021) · 2021
Earlier work this paper cites.
Cascadebert: Accelerating inference of pre-trained language models via calibrated complete models cascade
Li, L., Lin, Y., Chen, D., Ren, S., Li, P., Zhou, J., and Sun, X. (2021) · 2021
Earlier work this paper cites.
Pruning and quantization for deep neural network acceleration: A survey
Liang, T., Glossner, J., Wang, L., Shi, S., and Zhang, X. (2021) · 2021
Earlier work this paper cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N. A., and Lewis, M. (2021) · 2021
Earlier work this paper cites.
Step-unrolled denoising autoencoders for text generation
Savinov, N., Chung, J., Binkowski, M., Elsen, E., and Oord, A. v. d. (2021) · 2021
Earlier work this paper cites.
Consistent accelerated inference via confident adaptive transformers
Schuster, T., Fisch, A., Jaakkola, T., and Barzilay, R. (2021) · 2021
Earlier work this paper cites.
Lipschitz continuity guided knowledge distillation
Shang, Y., Duan, B., Zong, Z., Nie, L., and Yan, Y. (2021) · 2021
Earlier work this paper cites.
How many layers and why? An analysis of the model depth in transformers
Simoulin, A. and Crabbé, B. (2021) · 2021
Earlier work this paper cites.
Accelerating feedforward computation via parallel nonlinear equation solving
Song, Y., Meng, C., Liao, R., and Ermon, S. (2021) · 2021
Earlier work this paper cites.
LeeBERT: Learned early exit for BERT with cross-level optimization
Zhu, W. (2021) · 2021
Earlier work this paper cites.
A software-defined tensor streaming multiprocessor for large-scale machine learning
Abts, D., Kimmell, G., Ling, A., Kim, J., Boyd, M., Bitar, A., Parmar, S., Ahmed, I., DiCecco, R., Han, D., et al. (2022) · 2022
Cited alongside, same era.
Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale
Aminabadi, R. Y., Rajbhandari, S., Awan, A. A., Li, C., Li, D., Zheng, E., Ruwase, O., Smith, S., Zhang, M., Rasley, J., et al. (2022) · 2022
Cited alongside, same era.
Kerple: Kernelized relative positional embedding for length extrapolation
Chi, T.-C., Fan, T.-H., Ramadge, P. J., and Rudnicky, A. (2022) · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N. (2022) · 2022
Automoe: Heterogeneous mixture-of-experts with adaptive computation for efficient neural machine translation
Jawahar, G., Mukherjee, S., Liu, X., Kim, Y. J., Mageed, M. A., Laks Lakshmanan, V., Hassan, A., Bubeck, S., and Gao, J. (2023) · 2023
Later among the works it cites.
Early exit with disentangled representation and equiangular tight frame
Ji, Y., Wang, J., Li, J., Chen, Q., Chen, W., and Zhang, M. (2023) · 2023
Later among the works it cites.
Lion: Adversarial distillation of closed-source large language model
Jiang, Y., Chan, C., Chen, M., and Wang, W. (2023) · 2023
Later among the works it cites.
Flat: An optimized dataflow for mitigating attention bottlenecks
Kao, S.-C., Subramanian, S., Agrawal, G., Yazdanbakhsh, A., and Krishna, T. (2023) · 2023
Later among the works it cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I. (2023) · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C. (2022) · 2022
Cited alongside, same era.
Llm. int8 (): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L. (2022) · 2022
Cited alongside, same era.
Glam: Efficient scaling of language models with mixture-of-experts
Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., et al. (2022) · 2022
Cited alongside, same era.
A review of sparse expert models in deep learning
Fedus, W., Dean, J., and Zoph, B. (2022) · 2022
Cited alongside, same era.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D. (2022) · 2022
Cited alongside, same era.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Geva, M., Caciularu, A., Wang, K. R., and Goldberg, Y. (2022) · 2022
Cited alongside, same era.
A survey of quantization methods for efficient neural network inference
Gholami, A., Kim, S., Dong, Z., Yao, Z., Mahoney, M. W., and Keutzer, K. (2022) · 2022
Cited alongside, same era.
Ant: Exploiting adaptive numerical data type for low-bit deep neural network quantization
Guo, C., Zhang, C., Leng, J., Liu, Z., Yang, F., Liu, Y., Guo, M., and Zhu, Y. (2022a) · 2022
Cited alongside, same era.
Later among the works it cites.
Recallm: An adaptable memory mechanism with temporal understanding for large language models
Kynoch, B., Latapie, H., and van der Sluis, D. (2023) · 2023
Later among the works it cites.
Copy is all you need
Lan, T., Cai, D., Wang, Y., Huang, H., and Mao, X.-L. (2023) · 2023
Later among the works it cites.
Owq: Lessons learned from activation outliers for weight quantization in large language models
Lee, C., Jin, J., Kim, T., Kim, H., and Park, E. (2023) · 2023
Later among the works it cites.
Fast inference from transformers via speculative decoding
Leviathan, Y., Kalman, M., and Matias, Y. (2023) · 2023
Later among the works it cites.
Awq: Activation-aware weight quantization for llm compression and acceleration
Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S. (2023) · 2023
Later among the works it cites.
Llm-pruner: On the structural pruning of large language models
Ma, X., Fang, G., and Wang, X. (2023) · 2023
Later among the works it cites.
Ret-llm: Towards a general read-write memory for large language models
Modarressi, A., Imani, A., Fayyaz, M., and Schütze, H. (2023) · 2023
Later among the works it cites.
Landmark attention: Random-access infinite context length for transformers
Mohtashami, A. and Jaggi, M. (2023) · 2023
Later among the works it cites.
Pass: Parallel speculative sampling
Monea, G., Joulin, A., and Grave, E. (2023) · 2023
Later among the works it cites.
Skeleton-of-thought: Large language models can do parallel decoding
Ning, X., Lin, Z., Zhou, Z., Wang, Z., Yang, H., and Wang, Y. (2023) · 2023
Later among the works it cites.
Giraffe: Adventures in expanding context lengths in llms
Pal, A., Karkhanis, D., Roberts, M., Dooley, S., Sundararajan, A., and Naidu, S. (2023) · 2023
Later among the works it cites.
Pandya, K. and Holia, M. (2023) · 2023
Later among the works it cites.
Park, G., Park, B., Kim, M., Lee, S., Kim, J., Kwon, B., Kwon, S. J., Kim, B., Lee, Y., and Lee, D. (2023) · 2023
Later among the works it cites.
Yarn: Efficient context window extension of large language models
Peng, B., Quesnelle, J., Fan, H., and Shippole, E. (2023) · 2023
Later among the works it cites.
Fact: Ffn-attention co-optimized transformer architecture with eager correlation prediction
Qin, Y., Wang, Y., Deng, D., Zhao, Z., Yang, X., Liu, L., Wei, S., Hu, Y., and Yin, S. (2023) · 2023
Later among the works it cites.
Finding the sweet spot: Analysis and improvement of adaptive inference in low resource settings
Rotem, D., Hassid, M., Mamou, J., and Schwartz, R. (2023) · 2023
Later among the works it cites.
Promptmix: A class boundary augmentation method for large language model distillation
Sahu, G., Vechtomova, O., Bahdanau, D., and Laradji, I. H. (2023) · 2023
Later among the works it cites.
Accelerating transformer inference for translation via parallel decoding
Santilli, A., Severino, S., Postolache, E., Maiorca, V., Mancusi, M., Marin, R., and Rodolà, E. (2023) · 2023
Later among the works it cites.
The truth is in there: Improving reasoning in language models with layer-selective rank reduction
Sharma, P., Ash, J. T., and Misra, D. (2023) · 2023
Later among the works it cites.
Flexgen: High-throughput generative inference of large language models with a single GPU
Sheng, Y., Zheng, L., Yuan, B., Li, Z., Ryabinin, M., Chen, B., Liang, P., Ré, C., Stoica, I., and Zhang, C. (2023) · 2023
Later among the works it cites.
Powerinfer: Fast large language model serving with a consumer-grade gpu
Song, Y., Mi, Z., Xie, H., and Chen, H. (2023) · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al. (2023) · 2023
Later among the works it cites.
Tesla dojo technology: A guide to tesla’s configurable floating point formats & arithmetic
Tesla (2023) · 2023
Later among the works it cites.
Optimizing inference on large language models with nvidia tensorrt-llm, now publicly available
Vaidya, N., Oh, F., and Comly, N. (2023) · 2023
Later among the works it cites.
Scott: Self-consistent chain-of-thought distillation
Wang, P., Wang, Z., Li, Z., Gao, Y., Yin, B., and Ren, X. (2023) · 2023
Later among the works it cites.
Effective long-context scaling of foundation models
Xiong, W., Liu, J., Molybog, I., Zhang, H., Bhargava, P., Hou, R., Martin, L., Rungta, R., Sankararaman, K. A., Oguz, B., Khabsa, M., Fang, H., Mehdad, Y., Narang, S., Malik, K., Fan, A., Bhosale, S., Edunov, S., Lewis, M., Wang, S., and Ma, H. (2023) · 2023
Later among the works it cites.
Inference with reference: Lossless acceleration of large language models
Yang, N., Ge, T., Wang, L., Jiao, B., Jiang, D., Yang, L., Majumder, R., and Wei, F. (2023) · 2023
Later among the works it cites.
Edgemoe: Fast on-device inference of moe-based large language models
Yi, R., Guo, L., Wei, S., Zhou, A., Wang, S., and Xu, M. (2023) · 2023
Later among the works it cites.
Consistentee: A consistent and hardness-guided early exiting method for accelerating language models inference
Zeng, Z., Hong, Y., Dai, H., Zhuang, H., and Chen, C. (2023) · 2023
Later among the works it cites.
Draft & verify: Lossless large language model acceleration via self-speculative decoding
Zhang, J., Wang, J., Li, H., Shou, L., Chen, K., Chen, G., and Mehrotra, S. (2023) · 2023
Later among the works it cites.
A survey of large language models
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., et al. (2023) · 2023
Later among the works it cites.
Memorybank: Enhancing large language models with long-term memory
Zhong, W., Guo, L., Gao, Q., and Wang, Y. (2023) · 2023
Later among the works it cites.
Dimm-link: Enabling efficient inter-dimm communication for near-memory processing
Zhou, Z., Li, C., Yang, F., and Suny, G. (2023d) · 2023
Later among the works it cites.
A survey on model compression for large language models
Zhu, X., Li, J., Liu, Y., Ma, C., and Wang, W. (2023) · 2023
Later among the works it cites.
Unlimiformer: Long-range transformers with unlimited length input
Bertsch, A., Alon, U., Neubig, G., and Gormley, M. (2024) · 2024
Closest in time.
Medusa: Simple llm inference acceleration framework with multiple decoding heads
Cai, T., Li, Y., Geng, Z., Peng, H., Lee, J. D., Chen, D., and Dao, T. (2024) · 2024
Closest in time.
Mobilevlm v2: Faster and stronger baseline for vision language model
Chu, X., Qiao, L., Zhang, X., Xu, S., Wei, F., Yang, Y., Sun, X., Hu, Y., Lin, X., Zhang, B., et al. (2024) · 2024
Closest in time.
Tasks and tutorials using graphcore’s ipu with hugging face
Graphcore (2024) · 2024
Closest in time.
Deepspeed-fastgen: High-throughput text generation for llms via mii and deepspeed-inference
Holmes, C., Tanaka, M., Wyatt, M., Awan, A. A., Rasley, J., Rajbhandari, S., Aminabadi, R. Y., Qin, H., Bakhtiari, A., Kurilenko, L., et al. (2024) · 2024
Closest in time.
Kvquant: Towards 10 million context length llm inference with kv cache quantization
Hooper, C., Kim, S., Mohammadzadeh, H., Mahoney, M. W., Shao, Y. S., Keutzer, K., and Gholami, A. (2024) · 2024
Closest in time.
Billm: Pushing the limit of post-training quantization for llms
Huang, W., Liu, Y., Qin, H., Li, Y., Zhang, S., Liu, X., Magno, M., and Qi, X. (2024) · 2024
Closest in time.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al. (2024) · 2024
Closest in time.
Scalability limitations of processing-in-memory using real system evaluations
Jonatan, G., Cho, H., Son, H., Wu, X., Livesay, N., Mora, E., Shivdikar, K., Abellán, J. L., Joshi, A., Kaeli, D., et al. (2024) · 2024
Closest in time.
Eagle: Speculative sampling requires rethinking feature uncertainty
Li, Y., Wei, F., Zhang, C., and Zhang, H. (2024) · 2024
Closest in time.
Moe-llava: Mixture of experts for large vision-language models
Lin, B., Tang, Z., Ye, Y., Cui, J., Zhu, B., Jin, P., Zhang, J., Ning, M., and Yuan, L. (2024) · 2024
Closest in time.
Invariant test-time adaptation for vision-language model generalization
Ma, H., Zhu, Y., Zhang, C., Zhao, P., Wu, B., Huang, L.-K., Hu, Q., and Wu, B. (2024) · 2024
Closest in time.
Lightllm: A python-based llm inference and serving framework
ModelTC (2024) · 2024
Closest in time.
Dynamic memory compression: Retrofitting llms for accelerated inference
Nawrot, P., Łańcucki, A., Chochowski, M., Tarjan, D., and Ponti, E. M. (2024) · 2024
Closest in time.
Inside the nvidia hopper architecture
NVIDIA (2022) · 2024
Closest in time.
Pb-llm: Partially binarized large language models
Shang, Y., Yuan, Z., and Dong, Z. (2024) · 2024
Closest in time.
Our next-generation model: Gemini 1.5
Sundar Pichai, D. H. (2024) · 2024
Closest in time.
Small language model meets with reinforced vision vocabulary
Wei, H., Kong, L., Chen, J., Zhao, L., Ge, Z., Yu, E., Sun, J., Han, C., and Zhang, X. (2024) · 2024
Closest in time.
Fp6-llm: Efficiently serving large language models through fp6-centric algorithm-system co-design
Xia, H., Zheng, Z., Wu, X., Chen, S., Yao, Z., Youn, S., Bakhtiari, A., Wyatt, M., Zhuang, D., Zhou, Z., et al. (2024) · 2024
Closest in time.
A survey on knowledge distillation of large language models
Xu, X., Li, M., Tao, C., Shen, T., Cheng, R., Li, J., Xu, C., Tao, D., and Zhou, T. (2024) · 2024
Closest in time.
Wkvquant: Quantizing weight and key/value cache for large language models gains more
Yue, Y., Yuan, Z., Duanmu, H., Zhou, S., Wu, J., and Nie, L. (2024) · 2024
Closest in time.
Flightllm: Efficient large language model inference with a complete mapping flow on fpga
Zeng, S., Liu, J., Dai, G., Yang, X., Fu, T., Wang, H., Ma, W., Sun, H., Li, S., Huang, Z., et al. (2024) · 2024
Closest in time.
Retrieval-augmented generation for ai-generated content: A survey
Zhao, P., Zhang, H., Yu, Q., Wang, Z., Geng, Y., Fu, F., Yang, L., Zhang, W., and Cui, B. (2024) · 2024
Closest in time.
Tinyllava: A framework of small-scale large multimodal models
Zhou, B., Hu, Y., Weng, X., Jia, J., Luo, J., Liu, X., Wu, J., and Huang, L. (2024) · 2024
Closest in time.
Llava-phi: Efficient multi-modal assistant with small language model
Zhu, Y., Zhu, M., Liu, N., Ou, Z., Mou, X., and Tang, J. (2024) · 2024
Closest in time.