Fetching the paper…
Reading the bibliography…
The computational and memory challenges of large language models (LLMs) have sparked several optimization approaches towards their efficient implementation.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Edgebert: Sentence-level energy optimizations for latency-aware multi-task nlp inference
Tambe, T., Hooper, C., Pentecost, L., Jia, T., Yang, E.-Y., Donato, M., Sanh, V., Whatmough, P., Rush, A. M., Brooks, D., et al · 2021
Earlier work this paper cites.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H · 2021
Earlier work this paper cites.
An algorithm–hardware co-optimized framework for accelerating n: M sparse transformers
Fang, C., Zhou, A., and Wang, Z · 2022
Earlier work this paper cites.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Earlier work this paper cites.
Dynamic n: M fine-grained structured sparse attention mechanism
Chen, Z., Qu, Z., Quan, Y., Liu, L., Ding, Y., and Xie, Y · 2023
Earlier work this paper cites.
Llm-pruner: On the structural pruning of large language models
Ma, X., Fang, G., and Wang, X · 2023
Earlier work this paper cites.
Lingoqa: Visual question answering for autonomous driving
Marcu, A.-M., Chen, L., Hünermann, J., Karnsund, A., Hanotte, B., Chidananda, P., Nair, S., Badrinarayanan, V., Kendall, A., Shotton, J., and Sinavski, O · 2023
Earlier work this paper cites.
Fact: Ffn-attention co-optimized transformer architecture with eager correlation prediction
Qin, Y., Wang, Y., Deng, D., Zhao, Z., Yang, X., Liu, L., Wei, S., Hu, Y., and Yin, S · 2023
Cited alongside, same era.
22.9 a 12nm 18.1 tflops/w sparse transformer processor with entropy-based early exit, mixed-precision predication and fine-grained power management
Tambe, T., Zhang, J., Hooper, C., Jia, T., Whatmough, P. N., Zuckerman, J., Dos Santos, M. C., Loscalzo, E. J., Giri, D., Shepard, K., et al · 2023
Cited alongside, same era.
Cta: Hardware-software co-design for compressed token attention mechanism
Wang, H., Xu, H., Wang, Y., and Han, Y · 2023
Cited alongside, same era.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S · 2023
Cited alongside, same era.
A survey on evaluation of large language models
Chang, Y., Wang, X., Wang, J., Wu, Y., Yang, L., Zhu, K., Chen, H., Yi, X., Wang, C., Wang, Y., et al · 2024
Cited alongside, same era.
Large language models: A survey
Minaee, S., Mikolov, T., Nikzad, N., Chenaghlu, M., Socher, R., Amatriain, X., and Gao, J · 2024
Later among the works it cites.
Mobileaibench: Benchmarking llms and lmms for on-device use cases, 2024
Murthy, R., Yang, L., Tan, J., Awalgaonkar, T. M., Zhou, Y., Heinecke, S., Desai, S., Wu, J., Xu, R., Tan, S., Zhang, J., Liu, Z., Kokane, S., Liu, Z., Zhu, M., Wang, H., Xiong, C., and Savarese, S · 2024
Later among the works it cites.
An lpddr-based cxl-pnm platform for tco-efficient inference of transformer-based large language models
Park, S.-S., Kim, K., So, J., Jung, J., Lee, J., Woo, K., Kim, N., Lee, Y., Kim, H., Kwon, Y., et al · 2024
Later among the works it cites.
Mecla: Memory-compute-efficient llm accelerator with scaling sub-matrix partition
Qin, Y., Wang, Y., Zhao, Z., Yang, X., Zhou, Y., Wei, S., Hu, Y., and Yin, S · 2024
Later among the works it cites.
Llamaf: An efficient LLAMA2 architecture accelerator on embedded FPGAs
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ELSA: Exploiting Layer-wise N: M Sparsity for Vision Transformer Acceleration
Huang, N.-C., Chang, C.-C., Lin, W.-C., Taka, E., Marculescu, D., and Wu, K.-C · 2024
Cited alongside, same era.
AWQ: Activation-aware weight quantization for on-device LLM compression and acceleration
Lin, J., Tang, J., Tang, H., Yang, S., Chen, W.-M., Wang, W.-C., Xiao, G., Dang, X., Gan, C., and Han, S · 2024
Cited alongside, same era.
Flexllm: A system for co-serving large language model inference and parameter-efficient finetuning
Miao, X., Oliaro, G., Cheng, X., Wu, M., Unger, C., and Jia, Z · 2024
Cited alongside, same era.
URL {https://www.amd.com/en/products/accelerators/alveo.html}
AMD Alveo™ Adaptable Accelerator Cards
Cited in the paper.
URL {https://www.xilinx.com/products/boards-and-kits/ek-u1-zcu102-g.html}
Zynq UltraScale+ MPSoC ZCU102 Evaluation Kit, a
Cited in the paper.
URL {https://www.xilinx.com/products/boards-and-kits/zcu104.html}
Zynq UltraScale+ MPSoC ZCU104 Evaluation Kit, b
Cited in the paper.
Trex-reusing vision transformer’s attention for efficient xbar-based computing
Moitra, A., Bhattacharjee, A., Kim, Y., and Panda, P
Cited in the paper.
Xu, H., Li, Y., and Ji, S · 2024
Later among the works it cites.
Flightllm: Efficient large language model inference with a complete mapping flow on fpgas
Zeng, S., Liu, J., Dai, G., Yang, X., Fu, T., Wang, H., Ma, W., Sun, H., Li, S., Huang, Z., et al · 2024
Later among the works it cites.
LLMCompass: Enabling Efficient Hardware Design for Large Language Model Inference
Zhang, H., Ning, A., Prabhakar, R. B., and Wentzlaff, D · 2024
Later among the works it cites.
Alisa: Accelerating large language model inference via sparsity-aware kv caching
Zhao, Y., Wu, D., and Wang, J · 2024
Later among the works it cites.