Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated remarkable capabilities across various fields, from natural language understanding to text generation.
Implementing real-time robotic systems using chimera ii
David B Stewart, Donald E Schmitz, and Pradeep K Khosla · 1990
Earlier work this paper cites.
Moore’s law: past, present and future
Robert R Schaller · 1997
Earlier work this paper cites.
A parallel computing platform for real-time haptic interaction with deformable bodies
Ramin Mafi, Shahin Sirouspour, Behzad Mahdavikhah, Brian Moody, Kaveh Elizeh, Adam Kinsman, and Nicola Nicolici · 2009
Earlier work this paper cites.
CUDA programming: a developer’s guide to parallel computing with GPUs
Shane Cook · 2012
Earlier work this paper cites.
A case for exploiting subarray-level parallelism (salp) in dram
Yoongu Kim, Vivek Seshadri, Donghyuk Lee, Jamie Liu, and Onur Mutlu · 2012
Earlier work this paper cites.
The human brain project
Henry Markram · 2012
Earlier work this paper cites.
A cellular perspective on brain energy metabolism and functional imaging
Pierre J Magistretti and Igor Allaman · 2015
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Ultra-performance pascal gpu and nvlink interconnect
Denis Foley and John Danskin · 2017
Earlier work this paper cites.
Nvidia tesla v100 gpu architecture, 2017
NVIDIA Inc · 2017
Earlier work this paper cites.
Attention is all you need
A Vaswani · 2017
Earlier work this paper cites.
cublas: Basic linear algebra on nvidia gpus
NVIDIA · 2017
Earlier work this paper cites.
Cutlass: Cuda templates for linear algebra subroutines
NVIDIA · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Blockwise parallel decoding for deep autoregressive models
Mitchell Stern, Noam Shazeer, and Jakob Uszkoreit · 2018
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
M Lewis · 2019
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Earlier work this paper cites.
Adaptively sparse transformers
Gonçalo M Correia, Vlad Niculae, and André FT Martins · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2019
Earlier work this paper cites.
Electra: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning · 2020
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Earlier work this paper cites.
mt5: A massively multilingual pre-trained text-to-text transformer
L Xue · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Nvidia a100 tensor core gpu architecture, 2020
NVIDIA Inc · 2020
Earlier work this paper cites.
Amd cdna architecture, 2020
AMD Inc · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Earlier work this paper cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al · 2020
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan · 2020
Earlier work this paper cites.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Earlier work this paper cites.
Sparse sinkhorn attention
Yi Tay, Dara Bahri, Liu Yang, Donald Metzler, and Da-Cheng Juan · 2020
Earlier work this paper cites.
Intel®gaudi® ai accelerator first generation deep learning training and inference processor, 2020
Intel Inc · 2020
Earlier work this paper cites.
Glu variants improve transformer
Noam Shazeer · 2020
Earlier work this paper cites.
Think fast: A tensor streaming processor (tsp) for accelerating deep learning workloads
Dennis Abts, Jonathan Ross, Jonathan Sparling, Mark Wong-VanHaren, Max Baker, Tom Hawkins, Andrew Bell, John Thompson, Temesghen Kahsai, Garrin Kimmell, et al · 2020
Earlier work this paper cites.
A 4gbps dppm on-chip serial link based on pipelined vernier-tdc
Jinhao Li, Chong Qu, Fan Wu, and Jianfei Jiang · 2020
Earlier work this paper cites.
Glm: General language model pretraining with autoregressive blank infilling
Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang · 2021
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Earlier work this paper cites.
Wei Zeng, Xiaozhe Ren, Teng Su, Hui Wang, Yi Liao, Zhiwei Wang, Xin Jiang, ZhenZhang Yang, Kaisheng Wang, Xiaoda Zhang, et al · 2021
Earlier work this paper cites.
Ernie 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation
Yu Sun, Shuohuan Wang, Shikun Feng, Siyu Ding, Chao Pang, Junyuan Shang, Jiaxiang Liu, Xuyi Chen, Yanbin Zhao, Yuxiang Lu, et al · 2021
Earlier work this paper cites.
Multitask prompted training enables zero-shot task generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al · 2021
Earlier work this paper cites.
Shuohuan Wang, Yu Sun, Yang Xiang, Zhihua Wu, Siyu Ding, Weibao Gong, Shikun Feng, Junyuan Shang, Yanbin Zhao, Chao Pang, et al · 2021
Earlier work this paper cites.
Amd cdna 2 architecture, 2021
AMD Inc · 2021
Earlier work this paper cites.
Shuangfei Zhai, Walter Talbott, Nitish Srivastava, Chen Huang, Hanlin Goh, Ruixiang Zhang, and Josh Susskind · 2021
Earlier work this paper cites.
Efficiently modeling long sequences with structured state spaces
Albert Gu, Karan Goel, and Christopher Ré · 2021
Earlier work this paper cites.
Spatten: Efficient sparse attention architecture with cascade token and head pruning
Hanrui Wang, Zhekai Zhang, and Song Han · 2021
Earlier work this paper cites.
Hardware architecture and software stack for pim based on commercial dram technology: Industrial product
Sukhan Lee, Shin-haeng Kang, Jaehoon Lee, Hyeonsu Kim, Eojin Lee, Seungwoo Seo, Hosang Yoon, Seungwon Lee, Kyounghwan Lim, Hyunsung Shin, et al · 2021
Earlier work this paper cites.
A gain reconfigurable time difference amplifier with self-adaptive linearity control
Jinhao Li, Jianfei Jiang, Qin Wang, Naifeng Jing, Weiguang Sheng, and Guanghui He · 2021
Earlier work this paper cites.
What makes multi-modal learning better than single (provably)
Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen, Hang Zhao, and Longbo Huang · 2021
Earlier work this paper cites.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Earlier work this paper cites.
Glam: Efficient scaling of language models with mixture-of-experts
Nan Du, Yanping Huang, Andrew M Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, et al · 2022
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Earlier work this paper cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Earlier work this paper cites.
Lamda: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al · 2022
Earlier work this paper cites.
Nvidia h100 tensor core gpu architecture, 2022
NVIDIA Inc · 2022
Earlier work this paper cites.
Long range language modeling via gated state spaces
Harsh Mehta, Ankit Gupta, Ashok Cutkosky, and Behnam Neyshabur · 2022
Earlier work this paper cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh · 2022
Earlier work this paper cites.
Gunho Park, Baeseong Park, Minsub Kim, Sungjae Lee, Jeonghoon Kim, Beomseok Kwon, Se Jung Kwon, Byeongwook Kim, Youngjoo Lee, and Dongsoo Lee · 2022
Earlier work this paper cites.
Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer · 2022
Earlier work this paper cites.
Transpim: A memory-based acceleration via software-hardware co-design for transformer
Minxuan Zhou, Weihong Xu, Jaeyoung Kang, and Tajana Rosing · 2022
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Tri Dao et al · 2022
Earlier work this paper cites.
Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale
Reza Yazdani Aminabadi et al · 2022
Earlier work this paper cites.
Dfx: A low-latency multi-fpga appliance for accelerating transformer-based text generation
Seongmin Hong, Seungjae Moon, Junsoo Kim, Sungjae Lee, Minsub Kim, Dongsoo Lee, and Joo-Young Kim · 2022
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · 2022
Earlier work this paper cites.
A software-defined tensor streaming multiprocessor for large-scale machine learning
Dennis Abts, Garrin Kimmell, Andrew Ling, John Kim, Matt Boyd, Andrew Bitar, Sahil Parmar, Ibrahim Ahmed, Roberto DiCecco, David Han, et al · 2022
Earlier work this paper cites.
Compute express link®: An open industry-standard interconnect enabling heterogeneous data-centric computing
Debendra Das Sharma · 2022
Earlier work this paper cites.
A 1ynm 1.25 v 8gb, 16gb/s/pin gddr6-based accelerator-in-memory supporting 1tflops mac operation and various activation functions for deep-learning applications
Seongju Lee, Kyuyoung Kim, Sanghoon Oh, Joonhong Park, Gimoon Hong, Dongyoon Ka, Kyudong Hwang, Jeongje Park, Kyeongpil Kang, Jungyeon Kim, et al · 2022
Earlier work this paper cites.
System architecture and software stack for gddr6-aim
Yongkee Kwon, Kornijcuk Vladimir, Nahsung Kim, Woojae Shin, Jongsoon Won, Minkyu Lee, Hyunha Joo, Haerang Choi, Guhyun Kim, Byeongju An, et al · 2022
Earlier work this paper cites.
About tesla’s optimus robot brain, a robotics hardware and software computer architecture perspective, 2022
Acceleration Robotics · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Bloom: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al · 2023
Earlier work this paper cites.
Introducing llama: A foundational, 65-billion-parameter large language model
AI Meta · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, et al · 2023
Earlier work this paper cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Earlier work this paper cites.
The falcon series of open language models
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Mérouane Debbah, Étienne Goffinet, Daniel Hesslow, Julien Launay, Quentin Malartic, et al · 2023
Earlier work this paper cites.
Rwkv: Reinventing rnns for the transformer era
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Stella Biderman, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, et al · 2023
Earlier work this paper cites.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Earlier work this paper cites.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2023
Earlier work this paper cites.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al · 2023
Earlier work this paper cites.
Apple introduces m2 ultra, 2023
Apple Inc · 2023
Earlier work this paper cites.
Amd cdna 3 architecture, 2023
AMD Inc · 2023
Earlier work this paper cites.
Towards efficient generative large language model serving: A survey from algorithms to systems
Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Hongyi Jin, Tianqi Chen, and Zhihao Jia · 2023
Earlier work this paper cites.
A survey on model compression for large language models
Xunyu Zhu, Jian Li, Yong Liu, Can Ma, and Weiping Wang · 2023
Earlier work this paper cites.
Hyena hierarchy: Towards larger convolutional language models
Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher Ré · 2023
Earlier work this paper cites.
Efficient llm inference on cpus
Haihao Shen, Hanwen Chang, Bo Dong, Yu Luo, and Hengyu Meng · 2023
Earlier work this paper cites.
Spqr: A sparse-quantized representation for near-lossless llm weight compression
Tim Dettmers, Ruslan Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan Alistarh · 2023
Earlier work this paper cites.
Squeezellm: Dense-and-sparse quantization
Sehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong, Xiuyu Li, Sheng Shen, Michael W Mahoney, and Kurt Keutzer · 2023
Earlier work this paper cites.
Llm-mq: Mixed-precision quantization for efficient llm deployment
Shiyao Li, Xuefei Ning, Ke Hong, Tengxuan Liu, Luning Wang, Xiuhong Li, Kai Zhong, Guohao Dai, Huazhong Yang, and Yu Wang · 2023
Earlier work this paper cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han · 2023
Earlier work this paper cites.
Towards end-to-end 4-bit inference on generative large language models
Saleh Ashkboos, Ilia Markov, Elias Frantar, Tingxuan Zhong, Xincheng Wang, Jie Ren, Torsten Hoefler, and Dan Alistarh · 2023
Earlier work this paper cites.
Llm-fp4: 4-bit floating-point quantized transformers
Shih-yang Liu, Zechun Liu, Xijie Huang, Pingcheng Dong, and Kwang-Ting Cheng · 2023
Earlier work this paper cites.
A fast and flexible fpga-based accelerator for natural language processing neural networks
Suyeon Hur, Seongmin Na, Dongup Kwon, Joonsung Kim, Andrew Boutros, Eriko Nurvitadhi, and Jangwoo Kim · 2023
Earlier work this paper cites.
Olive: Accelerating large language models via hardware-friendly outlier-victim pair quantization
Cong Guo, Jiaming Tang, Weiming Hu, Jingwen Leng, Chen Zhang, Fan Yang, Yunxin Liu, Minyi Guo, and Yuhao Zhu · 2023
Earlier work this paper cites.
Llm-pruner: On the structural pruning of large language models
Xinyin Ma, Gongfan Fang, and Xinchao Wang · 2023
Earlier work this paper cites.
Sparsegpt: Massive language models can be accurately pruned in one-shot
Elias Frantar and Dan Alistarh · 2023
Earlier work this paper cites.
A simple and effective pruning approach for large language models
Mingjie Sun, Zhuang Liu, Anna Bair, and J. Zico Kolter · 2023
Earlier work this paper cites.
E-sparse: Boosting the large language model inference through entropy-based n: M sparsity
Yun Li, Lin Niu, Xipeng Zhang, Kai Liu, Jianchen Zhu, and Zhanhui Kang · 2023
Earlier work this paper cites.
Flash-llm: Enabling cost-effective and highly-efficient large generative model inference with unstructured sparsity
Haojun Xia, Zhen Zheng, Yuchao Li, Donglin Zhuang, Zhongzhu Zhou, Xiafei Qiu, Yong Li, Wei Lin, and Shuaiwen Leon Song · 2023
Earlier work this paper cites.
Deja vu: Contextual sparsity for efficient llms at inference time
Zichang Liu, Jue Wang, Tri Dao, Tianyi Zhou, Binhang Yuan, Zhao Song, Anshumali Shrivastava, Ce Zhang, Yuandong Tian, Christopher Re, et al · 2023
Earlier work this paper cites.
Faster causal attention over large sequences through sparse flash attention
Matteo Pagliardini, Daniele Paliotta, Martin Jaggi, and François Fleuret · 2023
Earlier work this paper cites.
Tf-mvp: Novel sparsity-aware transformer accelerator with mixed-length vector pruning
Eunji Yoo, Gunho Park, Jung Gyu Min, Se Jung Kwon, Baeseong Park, Dongsoo Lee, and Youngjoo Lee · 2023
Earlier work this paper cites.
Hardsea: Hybrid analog-reram clustering and digital-sram in-memory computing accelerator for dynamic sparse self-attention in transformer
Shiwei Liu, Chen Mu, Hao Jiang, Yunzhengmao Wang, Jinshan Zhang, Feng Lin, Keji Zhou, Qi Liu, and Chixiao Chen · 2023
Earlier work this paper cites.
Inference with reference: Lossless acceleration of large language models
Nan Yang, Tao Ge, Liang Wang, Binxing Jiao, Daxin Jiang, Linjun Yang, Rangan Majumder, and Furu Wei · 2023
Cited alongside, same era.
Draft & verify: Lossless large language model acceleration via self-speculative decoding
Jun Zhang, Jue Wang, Huan Li, Lidan Shou, Ke Chen, Gang Chen, and Sharad Mehrotra · 2023
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao · 2023
Cited alongside, same era.
Flash-decoding for long-context inference
Tri Dao et al · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon et al · 2023
Cited alongside, same era.
Eagle: Speculative sampling requires rethinking feature uncertainty
Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang · 2024
Closest in time.
Eagle-2: Faster inference of language models with dynamic draft trees
Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang · 2024
Closest in time.
Ouroboros: Speculative decoding with large model enhanced drafting
Weilin Zhao, Yuxiang Huang, Xu Han, Chaojun Xiao, Zhiyuan Liu, and Maosong Sun · 2024
Closest in time.
Sequoia: Scalable, robust, and hardware-aware speculative decoding
Zhuoming Chen, Avner May, Ruslan Svirschevski, Yuhsun Huang, Max Ryabinin, Zhihao Jia, and Beidi Chen · 2024
Closest in time.
Kangaroo: Lossless self-speculative decoding via double early exiting
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Openppl: A high-performance deep learning inference platform
Sensetime · 2023
Cited alongside, same era.
Optimizing inference on large language models with nvidia tensorrt-llm, now publicly available
Neal Vaidya et al · 2023
Cited alongside, same era.
Bytetransformer: A high-performance transformer boosted for variable-length inputs
Yujia Zhai et al · 2023
Cited alongside, same era.
Intel®gaudi® 2 ai accelerator high performance acceleration for genai and llms, 2023
Intel Inc · 2023
Cited alongside, same era.
Powerinfer: Fast large language model serving with a consumer-grade gpu
Yixin Song, Zeyu Mi, Haotong Xie, and Haibo Chen · 2023
Cited alongside, same era.
Unleashing the potential of pim: Accelerating large batched inference of transformer-based generative models
Jaewan Choi, Jaehyun Park, Kwanhee Kyung, Nam Sung Kim, and Jung Ho Ahn · 2023
Cited alongside, same era.
The era of generative artificial intelligence: In-memory computing perspective
Shin-haeng Kang, Sukhan Lee, and Kyomin Sohn · 2023
Cited alongside, same era.
Fangcheng Liu, Yehui Tang, Zhenhua Liu, Yunsheng Ni, Kai Han, and Yunhe Wang · 2024
Closest in time.
Layer skip: Enabling early exit inference and self-speculative decoding
Mostafa Elhoushi, Akshat Shrivastava, Diana Liskovich, Basil Hosmer, Bram Wasti, Liangzhen Lai, Anas Mahmoud, Bilge Acun, Saurabh Agarwal, Ahmed Roman, et al · 2024
Closest in time.
Not all layers of llms are necessary during inference
Siqi Fan, Xin Jiang, Xiang Li, Xuying Meng, Peng Han, Shuo Shang, Aixin Sun, Yequan Wang, and Zhongyuan Wang · 2024
Closest in time.
Raee: A training-free retrieval-augmented early exiting framework for efficient inference
Lianming Huang, Shangyu Wu, Yufei Cui, Ying Xiong, Xue Liu, Tei-Wei Kuo, Nan Guan, and Chun Jason Xue · 2024
Closest in time.
Mixture-of-depths: Dynamically allocating compute in transformer-based language models
David Raposo, Sam Ritter, Blake Richards, Timothy Lillicrap, Peter Conway Humphreys, and Adam Santoro · 2024
Closest in time.
Amusd: Asynchronous multi-device speculative decoding for llm acceleration
Bradley McDanel · 2024
Closest in time.
20.5 c-transformer: A 2.6-18.1 μ \mu j/token homogeneous dnn-transformer/spiking-transformer processor with big-little network and implicit weight generation for large language models
Sangyeob Kim, Sangjin Kim, Wooyoung Jo, Soyeon Kim, Seongyon Hong, and Hoi-Jun Yoo · 2024
Closest in time.
Specpim: Accelerating speculative inference on pim-enabled system via architecture-dataflow co-exploration
Cong Li, Zhe Zhou, Size Zheng, Jiaxi Zhang, Yun Liang, and Guangyu Sun · 2024
Closest in time.
Flashdecoding++: Faster large language model inference on gpus, 2024
Ke Hong et al · 2024
Closest in time.
Lpu: A latency-optimized and highly scalable processor for large language model inference
Seungjae Moon, Jung-Hoon Kim, Junsoo Kim, Seongmin Hong, Junseo Cha, Minsu Kim, Sukbin Lim, Gyubin Choi, Dongjin Seo, Jongho Kim, et al · 2024
Closest in time.
Consmax: Hardware-friendly alternative softmax with learnable parameters
Shiwei Liu, Guanchen Tao, Yifei Zou, Derek Chow, Zichen Fan, Kauna Lei, Bangfei Pan, Dennis Sylvester, Gregory Kielian, and Mehdi Saligane · 2024
Closest in time.
Marca: Mamba accelerator with reconfigurable architecture
Jinhao Li, Shan Huang, Jiaming Xu, Jun Liu, Li Ding, Ningyi Xu, and Guohao Dai · 2024
Closest in time.
Tcp: A tensor contraction processor for ai workloads industrial product
Hanjoon Kim, Younggeun Choi, Junyoung Park, Byeongwook Bae, Hyunmin Jeong, Sang Min Lee, Jeseung Yeon, Minho Kim, Changjae Park, Boncheol Gu, et al · 2024
Closest in time.
Intel gaudi 3 ai accelerator: Architected for gen ai training and inference
Roman Kaplan · 2024
Closest in time.
Balanced data placement for gemv acceleration with processing-in-memory
Mohamed Assem Ibrahim, Mahzabeen Islam, and Shaizeen Aga · 2024
Closest in time.
Rongqing Cong, Wenyang He, Mingxuan Li, Bangning Luo, Zebin Yang, Yuchao Yang, Ru Huang, and Bonan Yan · 2024
Closest in time.
Pim gpt a hybrid process in memory accelerator for autoregressive transformers
Yuting Wu, Ziyu Wang, and Wei D Lu · 2024
Closest in time.
Wontak Han, Hyunjun Cho, Donghyuk Kim, and Joo-Young Kim · 2024
Closest in time.
Pipepim: Maximizing computing unit utilization in ml-oriented digital pim by pipelining and dual buffering
Taeyang Jeong and Eui-Young Chung · 2024
Closest in time.
Exploiting intel® advanced matrix extensions (amx) for large language model inference
Hyungyo Kim, Gaohan Ye, Nachuan Wang, Amir Yazdanbakhsh, and Nam Sung Kim · 2024
Closest in time.
Powerinfer-2: Fast large language model inference on a smartphone
Zhenliang Xue, Yixin Song, Zeyu Mi, Le Chen, Yubin Xia, and Haibo Chen · 2024
Closest in time.
Twinpilots: A new computing paradigm for gpu-cpu parallel llm inference
Chengye Yu, Tianyu Wang, Zili Shao, Linjie Zhu, Xu Zhou, and Song Jiang · 2024
Closest in time.
Glitches: Gpu-fpga llm inference through a collaborative heterogeneous system
Fan Yang, Xinhao Yang, Hongyi Wang, Zehao Wang, Zhenhua Zhu, Shulin Zeng, and Yu Wang · 2024
Closest in time.
Neupims: Npu-pim heterogeneous acceleration for batched llm inferencing
Guseul Heo, Sangyeop Lee, Jaehong Cho, Hyunmin Choi, Sanghyeon Lee, Hyungkyu Ham, Gwangsun Kim, Divya Mahajan, and Jongse Park · 2024
Closest in time.
Ianus: Integrated accelerator based on npu-pim unified memory system
Minseok Seo, Xuan Truong Nguyen, Seok Joong Hwang, Yongkee Kwon, Guhyun Kim, Chanwook Park, Ilkon Kim, Jaehan Park, Jeongbin Kim, Woojae Shin, et al · 2024
Closest in time.
Monde: Mixture of near-data experts for large-scale sparse models
Taehyun Kim, Kwanseok Choi, Youngmock Cho, Jaehoon Cho, Hyuk-Jae Lee, and Jaewoong Sim · 2024
Closest in time.
Attacc! unleashing the power of pim for batched transformer-based generative model inference
Jaehyun Park, Jaewan Choi, Kwanhee Kyung, Michael Jaemin Kim, Yongsuk Kwon, Nam Sung Kim, and Jung Ho Ahn · 2024
Closest in time.
The breakthrough memory solutions for improved performance on llm inference
Byeongho Kim, Sanghoon Cha, Sangsoo Park, Jieun Lee, Sukhan Lee, Shin-haeng Kang, Jinin So, Kyungsoo Kim, Jin Jung, Jong-Geon Lee, et al · 2024
Closest in time.
H3d-transformer: A heterogeneous 3d (h3d) computing platform for transformer model acceleration on edge devices
Yandong Luo and Shimeng Yu · 2024
Closest in time.
An lpddr-based cxl-pnm platform for tco-efficient inference of transformer-based large language models
Sang-Soo Park, KyungSoo Kim, Jinin So, Jin Jung, Jonggeon Lee, Kyoungwan Woo, Nayeon Kim, Younghyun Lee, Hyungyo Kim, Yongsuk Kwon, et al · 2024
Closest in time.
Sk hynix ai-specific computing memory solution: From aim device to heterogeneous aimx-xpu system for comprehensive llm inference
Guhyun Kim, Jinkwon Kim, Nahsung Kim, Woojae Shin, Jongsoon Won, Hyunha Joo, Haerang Choi, Byeongju An, Gyeongcheol Shin, Dayeon Yun, et al · 2024
Closest in time.
Cambricon-llm: A chiplet-based hybrid architecture for on-device inference of 70b llm
Zhongkai Yu, Shengwen Liang, Tianyun Ma, Yunke Cai, Ziyuan Nan, Di Huang, Xinkai Song, Yifan Hao, Jie Zhang, Tian Zhi, et al · 2024
Closest in time.
Pim-ai: A novel architecture for high-efficiency llm inference
Cristobal Ortega, Yann Falevoz, and Renaud Ayrignac · 2024
Closest in time.
Distributed inference performance optimization for llms on cpus
Pujiang He, Shan Zhou, Changqing Li, Wenhuan Huang, Weifei Yu, Duyi Wang, Chen Meng, and Sheng Gui · 2024
Closest in time.
Inference performance optimization for large language models on cpus
Pujiang He, Shan Zhou, Wenhuan Huang, Changqing Li, Duyi Wang, Bin Guo, Chen Meng, Sheng Gui, Weifei Yu, and Yi Xie · 2024
Closest in time.
An agile framework for efficient llm accelerator development and model inference
Lvcheng Chen, Ying Wu, Chenyi Wen, Shizhang Wang, Li Zhang, Bei Yu, Qi Sun, and Cheng Zhuo · 2024
Closest in time.
Raspberry pi 5, 2024
Raspberry Pi · 2024
Closest in time.
Groundbreaking gemma 7b performance running on the groq lpu™ inference engine, 2024
Groq Inc · 2024
Closest in time.
Will we run out of data? limits of llm scaling based on human-generated data, 2024
Pablo Villalobos, Anson Ho, Jaime Sevilla, Tamay Besiroglu, Lennart Heim, and Marius Hobbhahn · 2024
Closest in time.
Mm-llms: Recent advances in multimodal large language models
Duzhen Zhang, Yahan Yu, Jiahua Dong, Chenxing Li, Dan Su, Chenhui Chu, and Dong Yu · 2024
Closest in time.
Sequence to sequence learning with neural networks at neurips 2024, 2024
Ilya Sutskever · 2024
Closest in time.
Introducing openai o1, 2024
OpenAI · 2024
Closest in time.
Openai o1-mini, September 2024
OpenAI · 2024
Closest in time.
Shibo Hao, Yi Gu, Haotian Luo, Tianyang Liu, Xiyan Shao, Xinyuan Wang, Shuhua Xie, Haodi Ma, Adithya Samavedhi, Qiyue Gao, et al · 2024
Closest in time.
Nvidia jetson orin, 2024
NVIDIA · 2024
Closest in time.
Le Chen, Dahu Feng, Erhu Feng, Rong Zhao, Yingrui Wang, Yubin Xia, Haibo Chen, and Pinjie Xu · 2025
Closest in time.
Deca: A near-core llm decompression accelerator supporting out-of-order invocation
Gerasimos Gerogiannis, Stijn Eyerman, Evangelos Georganas, Wim Heirman, and Josep Torrellas · 2025
Closest in time.
Accurate kv cache quantization with outlier tokens tracing
Yi Su, Yuechi Zhou, Quantong Qiu, Juntao Li, Qingrong Xia, Ping Li, Xinyu Duan, Zhefeng Wang, and Min Zhang · 2025
Closest in time.
Efficient arbitrary precision acceleration for large language models on gpu tensor cores
Shaobo Ma, Chao Fang, Haikuo Shao, and Zhongfeng Wang · 2025
Closest in time.
Edgellm: A highly efficient cpu-fpga heterogeneous edge accelerator for large language models
Mingqiang Huang, Ao Shen, Kai Li, Haoxiang Peng, Boyu Li, Yupeng Su, and Hao Yu · 2025
Closest in time.
On-device qwen2. 5: Efficient llm inference with model compression and hardware acceleration
Maoyang Xiang, Ramesh Fernando, and Bo Wang · 2025
Closest in time.
Lightmamba: Efficient mamba acceleration on fpga with quantization and hardware co-design
Renjie Wei, Songqiang Xu, Linfeng Zhong, Zebin Yang, Qingyu Guo, Yuan Wang, Runsheng Wang, and Meng Li · 2025
Closest in time.
Tellme: An energy-efficient ternary llm accelerator for prefilling and decoding on edge fpgas
Ye Qiao, Zhiheng Cheng, Yifan Zhang, Yian Wang, and Sitao Huang · 2025
Closest in time.
Tereffic: Highly efficient ternary llm inference on fpga
Chenyang Yin, Zhenyu Bai, Pranav Venkatram, Shivam Aggarwal, Zhaoying Li, and Tulika Mitra · 2025
Closest in time.
Jindong Li, Tenglong Li, Guobin Shen, Dongcheng Zhao, Qian Zhang, and Yi Zeng · 2025
Closest in time.
Meadow: Memory-efficient dataflow and data packing for low power edge llms
Abhishek Moitra, Arkapravo Ghosh, Shrey Agarwal, Aporva Amarnath, Karthik Swaminathan, and Priyadarshini Panda · 2025
Closest in time.
A tensor-train decomposition based compression of llms on group vector systolic accelerator
Sixiao Huang, Tintin Wang, Ang Li, Ao Shen, Kai Li, Keyao Jiang, Mingqiang Huang, and Hao Yu · 2025
Closest in time.
Looplynx: A scalable dataflow architecture for efficient llm inference
Jianing Zheng and Gang Chen · 2025
Closest in time.
Accllm: Accelerating long-context llm inference via algorithm-hardware co-design
Yanbiao Liang, Huihong Shi, Haikuo Shao, and Zhongfeng Wang · 2025
Closest in time.
Adaptive two-range quantization and hardware co-design for large language model acceleration
Siqi Cai, Gang Wang, Wenjie Li, Dongxu Lyu, and Guanghui He · 2025
Closest in time.
Fineq: Software-hardware co-design for low-bit fine-grained mixed-precision quantization of llms
Xilong Xie, Liang Wang, Limin Xiao, Meng Han, Lin Sun, Shuai Zheng, and Xiangrong Xu · 2025
Closest in time.
Bitmod: Bit-serial mixture-of-datatype llm acceleration
Yuzong Chen, Ahmed F AbouElhamayed, Xilai Dai, Yang Wang, Marta Andronic, George A Constantinides, and Mohamed S Abdelfattah · 2025
Closest in time.
Ofq-llm: Outlier-flexing quantization for efficient low-bit large language model acceleration
Gang Wang, Siqi Cai, Wenjie Li, Dongxu Lyu, and Guanghui He · 2025
Closest in time.
M-ant: Efficient low-bit group quantization for llms via mathematically adaptive numerical type
Weiming Hu, Haoyan Zhang, Cong Guo, Yu Feng, Renyang Guan, Zhendong Hua, Zihan Liu, Yue Guan, Minyi Guo, and Jingwen Leng · 2025
Closest in time.
Anda: Unlocking efficient llm inference with a variable-length grouped activation data format
Chao Fang, Man Shi, Robin Geens, Arne Symons, Zhongfeng Wang, and Marian Verhelst · 2025
Closest in time.
Ecco: Improving memory bandwidth and capacity for llms via entropy-aware cache compression
Feng Cheng, Cong Guo, Chiyue Wei, Junyao Zhang, Changchun Zhou, Edward Hanson, Jiaqi Zhang, Xiaoxiao Liu, Hai Li, Yiran Chen, et al · 2025
Closest in time.
Integer unit-based outlier-aware llm accelerator preserving numerical accuracy of fp-fp gemm
Jehun Lee and Jae-Joon Kim · 2025
Closest in time.
Coleman Hooper, Charbel Sakr, Ben Keller, Rangharajan Venkatesan, Kurt Keutzer, Sophia Shao, and Brucek Khailany · 2025
Closest in time.
Lightrot: A light-weighted rotation scheme and architecture for accurate low-bit large language model inference
Sangjin Kim, Yuseon Choi, Jungjun Oh, Byeongcheol Kim, and Hoi-Jun Yoo · 2025
Closest in time.
Softmap: Software-hardware co-design for integer-only softmax on associative processors
Mariam Rakka, Jinhao Li, Guohao Dai, Ahmed Eltawil, Mohammed E Fouda, and Fadi Kurdahi · 2025
Closest in time.
Mvdram: Enabling gemv execution in unmodified dram for low-bit llm acceleration
Tatsuya Kubo, Daichi Tokuda, Tomoya Nagatani, Masayuki Usui, Lei Qu, Ting Cao, and Shinya Takamaeda-Yamazaki · 2025
Closest in time.
Pim-llm: A high-throughput hybrid pim architecture for 1-bit llms
Jinendra Malekar, Peyton Chandarana, Md Hasibul Amin, Mohammed E Elbtity, and Ramtin Zand · 2025
Closest in time.
Roma: a read-only-memory-based accelerator for qlora-based on-device llm
Wenqiang Wang, Yijia Zhang, Zikai Zhang, Guanting Huo, Hao Liang, Shijie Cao, and Ningyi Xu · 2025
Closest in time.
Sparamx: Accelerating compressed llms token generation on amx-powered cpus
Ahmed F AbouElhamayed, Jordan Dotzel, Yash Akhauri, Chi-Chih Chang, Sameh Gobriel, J Pablo Muñoz, Vui Seng Chua, Nilesh Jain, and Mohamed S Abdelfattah · 2025
Closest in time.
Efficient unstructured pruning of mamba state-space models for resource-constrained environments
Ibne Farabi Shihab, Sanjeda Akter, and Anuj Sharma · 2025
Closest in time.
Spinfer: Leveraging low-level sparsity for efficient large language model inference on gpus
Ruibo Fan, Xiangrui Yu, Peijie Dong, Zeyu Li, Gu Gong, Qiang Wang, Wei Wang, and Xiaowen Chu · 2025
Closest in time.
Sola: Leveraging soft activation sparsity and low-rank decomposition for large language model compression
Xinhao Huang, You-Liang Huang, and Zeyi Wen · 2025
Closest in time.
R-sparse: Rank-aware activation sparsity for efficient llm inference
Zhenyu Zhang, Zechun Liu, Yuandong Tian, Harshit Khaitan, Zhangyang Wang, and Steven Li · 2025
Closest in time.
Akshat Ramachandran, Souvik Kundu, Arnab Raha, Shamik Kundu, Deepak K Mathaikutty, and Tushar Krishna · 2025
Closest in time.
Ml-specqd: Multi-level speculative decoding with quantized drafts
Evangelos Georganas, Dhiraj Kalamkar, Alexander Kozlov, and Alexander Heinecke · 2025
Closest in time.
Specee: Accelerating large language model inference with speculative early exiting
Jiaming Xu, Jiayi Pan, Yongkang Zhou, Siming Chen, Jinhao Li, Yaoxiu Lian, Junyi Wu, and Guohao Dai · 2025
Closest in time.
Pipespec: Breaking stage dependencies in hierarchical llm decoding
Bradley McDanel, Sai Qian Zhang, Yunhai Hu, and Zining Liu · 2025
Closest in time.
Adaptive draft-verification for efficient large language model decoding
Xukun Liu, Bowen Lei, Ruqi Zhang, and Dongkuan DK Xu · 2025
Closest in time.
Spin: Accelerating large language model inference with heterogeneous speculative models
Fahao Chen, Peng Li, Tom H Luan, Zhou Su, and Jing Deng · 2025
Closest in time.
Pard: Accelerating llm inference with low-cost parallel draft model adaptation
Zihao An, Huajun Bai, Ziqiong Liu, Dong Li, and Emad Barsoum · 2025
Closest in time.
Judge decoding: Faster speculative sampling requires going beyond model alignment
Gregor Bachmann, Sotiris Anagnostidis, Albert Pumarola, Markos Georgopoulos, Artsiom Sanakoyeu, Yuming Du, Edgar Schönfeld, Ali Thabet, and Jonas Kohler · 2025
Closest in time.
Falcon: Faster and parallel inference of large language models through enhanced semi-autoregressive drafting and custom-designed decoding tree
Xiangxiang Gao, Weisheng Xie, Yiwei Xiang, and Feng Ji · 2025
Closest in time.
V-seek: Accelerating llm reasoning on open-hardware server-class risc-v platforms, 2025
Javier J. Poveda Rodrigo, Mohamed Amine Ahmdi, Alessio Burrello, Daniele Jahier Pagliari, and Luca Benini · 2025
Closest in time.
Flashformer: Whole-model kernels for efficient low-batch inference
Aniruddha Nrusimha, William Brandon, Mayank Mishra, Yikang Shen, Rameswar Panda, Jonathan Ragan-Kelley, and Yoon Kim · 2025
Closest in time.
Accelerating llm inference throughput via asynchronous kv cache prefetching
Yanhao Dong, Yubo Miao, Weinan Li, Xiao Zheng, Chao Wang, and Feng Lyu · 2025
Closest in time.
Haan: A holistic approach for accelerating normalization operations in large language models
Tianfan Peng, Tianhua Xia, Jiajun Qin, and Sai Qian Zhang · 2025
Closest in time.
Picachu: Plug-in cgra handling upcoming nonlinear operations in llms
Jiajun Qin, Tianhua Xia, Cheng Tan, Jeff Zhang, and Sai Qian Zhang · 2025
Closest in time.
Waferllm: A wafer-scale llm inference system
Congjie He, Yeqi Huang, Pei Mu, Ziming Miao, Jilong Xue, Lingxiao Ma, Fan Yang, and Luo Mai · 2025
Closest in time.
Characterizing and optimizing llm inference workloads on cpu-gpu coupled architectures
Prabhu Vellaisamy, Thomas Labonte, Sourav Chakraborty, Matt Turner, Samantika Sury, and John Paul Shen · 2025
Closest in time.
Prima. cpp: Speeding up 70b-scale llm inference on low-resource everyday home clusters
Zonghang Li, Tao Li, Wenjiao Feng, Mohsen Guizani, and Hongfang Yu · 2025
Closest in time.
Tightllm: Maximizing throughput for llm inference via adaptive offloading policy
Yitao Hu, Xiulong Liu, Guotao Yang, Linxuan Li, Kai Zeng, Zhixin Zhao, Sheng Chen, Laiping Zhao, Wenxin Li, and Keqiu Li · 2025
Closest in time.
Hpu: High-bandwidth processing unit for scalable, cost-effective llm inference via gpu co-processing
Myunghyun Rhee, Joonseop Sim, Taeyoung Ahn, Seungyong Lee, Daegun Yoon, Euiseok Kim, Kyoung Park, Youngpyo Joo, and Hosik Kim · 2025
Closest in time.
Paise: Pim-accelerated inference scheduling engine for transformer-based llm
Hyojung Lee, Daehyeon Baek, Jimyoung Son, Jieun Choi, Kihyo Moon, and Minsung Jang · 2025
Closest in time.
Pyramid: Accelerating llm inference with cross-level processing-in-memory
Liang Yan, Xiaoyang Lu, Xiaoming Chen, Yinhe Han, and Xian-He Sun · 2025
Closest in time.
Make llm inference affordable to everyone: Augmenting gpu memory with ndp-dimm
Lian Liu, Shixin Zhao, Bing Li, Haimeng Ren, Zhaohui Xu, Mengdi Wang, Xiaowei Li, Yinhe Han, and Ying Wang · 2025
Closest in time.
Opencompass multi-modal leaderboard, 2025
OpenCompass · 2025
Closest in time.
Fine-grained modeling and evaluation of 3d power distribution networks in logic-on-memory stacking architecture
Ang Li, Pengyu Liu, Zizheng Dong, Jinhao Li, Jianfei Jiang, Weijia Zhu, Naifeng Jing, and Qin Wang · 2025
Closest in time.