Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have achieved remarkable success across various domains, yet deploying them on mobile devices remains an arduous challenge due to their extensive computational and memory demands.
An improved equivalence algorithm
Bernard A Galler and Michael J Fisher · 1964
Earlier work this paper cites.
A data structure for manipulating priority queues
Jean Vuillemin · 1978
Earlier work this paper cites.
Computers and intractability: a guide to the theory of np-completeness (michael r. garey and david s. johnson)
Juris Hartmanis · 1982
Earlier work this paper cites.
Comparing biases for minimal network construction with back-propagation
Stephen Hanson and Lorien Pratt · 1988
Earlier work this paper cites.
Introduction to the theory of computation
Michael Sipser · 1996
Earlier work this paper cites.
Model selection and estimation in regression with grouped variables
Ming Yuan and Yi Lin · 2006
Earlier work this paper cites.
Optimization of sparse matrix-vector multiplication on emerging multicore platforms
Samuel Williams, Leonid Oliker, Richard Vuduc, John Shalf, Katherine Yelick, and James Demmel · 2007
Earlier work this paper cites.
Efficient sparse matrix-vector multiplication on cuda
Nathan Bell and Michael Garland · 2008
Earlier work this paper cites.
Parallel sparse matrix-vector and matrix-transpose-vector multiplication using compressed sparse blocks
Aydin Buluç, Jeremy T Fineman, Matteo Frigo, John R Gilbert, and Charles E Leiserson · 2009
Earlier work this paper cites.
Calculate iops in a storage array
Scott Lowe · 2010
Earlier work this paper cites.
Cusparse library
Maxim Naumov, L Chien, Philippe Vandermersch, and Ujval Kapasi · 2010
Earlier work this paper cites.
Low-rank approximations for conditional feedforward computation in deep neural networks
Andrew Davis and Itamar Arel · 2013
Earlier work this paper cites.
Conditional computation in neural networks for faster models
Emmanuel Bengio, Pierre-Luc Bacon, Joelle Pineau, and Doina Precup · 2015
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally · 2015
Earlier work this paper cites.
Dynamic network surgery for efficient dnns
Yiwen Guo, Anbang Yao, and Yurong Chen · 2016
Earlier work this paper cites.
Fast convnets using group-wise brain damage
Vadim Lebedev and Victor Lempitsky · 2016
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
A survey of model compression and acceleration for deep neural networks
Yu Cheng, Duo Wang, Pan Zhou, and Tao Zhang · 2017
Earlier work this paper cites.
Deep learning using rectified linear units (relu)
AF Agarap · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin · 2018
Earlier work this paper cites.
Soft filter pruning for accelerating deep convolutional neural networks
Yang He, Guoliang Kang, Xuanyi Dong, Yanwei Fu, and Yi Yang · 2018
Earlier work this paper cites.
Data-driven sparse structure selection for deep neural networks
Zehao Huang and Naiyan Wang · 2018
Earlier work this paper cites.
Swag: A large-scale adversarial dataset for grounded commonsense inference
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi · 2018
Earlier work this paper cites.
Openwebtext corpus
Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex · 2019
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need
Noam Shazeer · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Sparse GPU kernels for deep learning
Trevor Gale, Matei Zaharia, Cliff Young, and Erich Elsen · 2020
Earlier work this paper cites.
Sparse gpu kernels for deep learning
Trevor Gale, Matei Zaharia, Cliff Young, and Erich Elsen · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Inducing and exploiting activation sparsity for fast inference on deep neural networks
Mark Kurtz, Justin Kopinsky, Rati Gelashvili, Alexander Matveev, John Carr, Michael Goin, William Leiserson, Sage Moore, Nir Shavit, and Dan Alistarh · 2020
Earlier work this paper cites.
Pruning algorithms to accelerate convolutional neural networks for edge applications: A survey
Jiayi Liu, Samarth Tripathi, Unmesh Kurup, and Mohak Shah · 2020
Earlier work this paper cites.
Movement pruning: Adaptive sparsity by fine-tuning
Victor Sanh, Thomas Wolf, and Alexander Rush · 2020
Earlier work this paper cites.
Sparsert: Accelerating unstructured sparsity on gpus for deep learning inference
Ziheng Wang · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Cited alongside, same era.
Pruning and quantization for deep neural network acceleration: A survey
Tailin Liang, John Glossner, Lei Wang, Shaobo Shi, and Xiaotong Zhang · 2021
Cited alongside, same era.
Rethinking network pruning–under the pre-train and fine-tune paradigm
Dongkuan Xu, Ian EH Yen, Jinxi Zhao, and Zhibin Xiao · 2021
Cited alongside, same era.
Moefication: Transformer feed-forward layers are mixtures of experts
Zhengyan Zhang, Yankai Lin, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou · 2021
Cited alongside, same era.
Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer · 2022
Cited alongside, same era.
Phi-3 technical report: A highly capable language model locally on your phone
Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl, et al · 2024
Closest in time.
Quip: 2-bit quantization of large language models with guarantees
Jerry Chee, Yaohui Cai, Volodymyr Kuleshov, and Christopher M De Sa · 2024
Closest in time.
Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models
Damai Dai, Chengqi Deng, Chenggang Zhao, RX Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Yu Wu, et al · 2024
Closest in time.
ggerganov/llama.cpp: Port of facebook’s llama model in c/c++
Georgi Gerganov · 2024
Closest in time.
Chess: Optimizing llm inference via channel-wise thresholding and selective sparsification
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2022
Cited alongside, same era.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh · 2022
Cited alongside, same era.
A survey of quantization methods for efficient neural network inference
Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W Mahoney, and Kurt Keutzer · 2022
Cited alongside, same era.
The lazy neuron phenomenon: On emergence of activation sparsity in transformers
Zonglin Li, Chong You, Srinadh Bhojanapalli, Daliang Li, Ankit Singh Rawat, Sashank J Reddi, Ke Ye, Felix Chern, Felix Yu, Ruiqi Guo, et al · 2022
Cited alongside, same era.
ChatGPT: Get instant answers, find creative inspiration, learn something new
OpenAI · 2022
Cited alongside, same era.
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Zhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu, Conglong Li, and Yuxiong He · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Cited alongside, same era.
Junhui He, Shangyu Wu, Weidong Wen, Chun Jason Xue, and Qingan Li · 2024
Closest in time.
Jedec announces publication of universal flash storage (ufs) standard
JEDEC · 2024
Closest in time.
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al · 2024
Closest in time.
Realization of low in-band harmonic for compact 6–18-ghz t/r module under tx-mode operation
Jinming Lai, Zhiyou Li, Chaojie Wang, Hailong Wang, and Xiaohua Ma · 2024
Closest in time.
Cats: Contextually-aware thresholding for sparsity in large language models
Donghyun Lee, Je-Yong Lee, Genghan Zhang, Mo Tiwari, and Azalia Mirhoseini · 2024
Closest in time.
Breaking relu barrier: Generalized moefication for dense pretrained models
Jaeseong Lee, Seung-won Hwang, Wonpyo Park, and Mingi Ji · 2024
Closest in time.
Personal llm agents: Insights and survey about the capability, efficiency and security
Yuanchun Li, Hao Wen, Weijun Wang, Xiangyu Li, Yizhen Yuan, Guohong Liu, Jiacheng Liu, Wenxing Xu, Xiang Wang, Yi Sun, et al · 2024
Closest in time.
Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model
Aixin Liu, Bei Feng, Bin Wang, Bingxuan Wang, Bo Liu, Chenggang Zhao, Chengqi Dengr, Chong Ruan, Damai Dai, Daya Guo, et al · 2024
Closest in time.
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al · 2024
Closest in time.
A contemporary overview: Trends and applications of large language models on mobile devices
Lianjun Liu, Hongli An, Pengxuan Chen, and Longxiang Ye · 2024
Closest in time.
Sparsing law: Towards large language models with greater activation sparsity
Yuqi Luo, Chenyang Song, Xu Han, Yingfa Chen, Chaojun Xiao, Zhiyuan Liu, and Maosong Sun · 2024
Closest in time.
Accelerating inference with sparsity using the nvidia ampere architecture and nvidia tensorrt, 2021
NVIDIA · 2024
Closest in time.
Prosparse: Introducing and enhancing intrinsic activation sparsity within large language models
Chenyang Song, Xu Han, Zhengyan Zhang, Shengding Hu, Xiyu Shi, Kuai Li, Chen Chen, Zhiyuan Liu, Guangli Li, Tao Yang, et al · 2024
Closest in time.
Achieving sparse activation in small language models
Jifeng Song, Kai Huang, Xiangyu Yin, Boyuan Yang, and Wei Gao · 2024
Closest in time.
Turbo sparse: Achieving llm sota performance with minimal activated parameters
Yixin Song, Haotong Xie, Zhengyan Zhang, Bo Wen, Li Ma, Zeyu Mi, and Haibo Chen · 2024
Closest in time.
Gpt-4 technical report, 2024
OpenAI Teams · 2024
Closest in time.
Mobillama: Towards accurate and lightweight fully transparent gpt, 2024
Omkar Thawakar, Ashmal Vayani, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer, Michael Felsberg, Timothy Baldwin, Eric P. Xing, and Fahad Shahbaz Khan · 2024
Closest in time.
Q-sparse: All large language models can be fully sparsely-activated
Hongyu Wang, Shuming Ma, Ruiping Wang, and Furu Wei · 2024
Closest in time.
Long exposure: Accelerating parameter-efficient fine-tuning for llms under shadowy sparsity
Tuowei Wang, Kun Li, Zixu Hao, Donglin Bai, Ju Ren, Yaoxue Zhang, Ting Cao, and Mao Yang · 2024
Closest in time.
Autodroid: Llm-powered task automation in android
Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby Jia-Jun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu · 2024
Closest in time.
Onebit: Towards extremely low-bit large language models
Yuzhuang Xu, Xu Han, Zonghan Yang, Shuo Wang, Qingfu Zhu, Zhiyuan Liu, Weidong Liu, and Wanxiang Che · 2024
Closest in time.
Powerinfer-2: Fast large language model inference on a smartphone
Zhenliang Xue, Yixin Song, Zeyu Mi, Le Chen, Yubin Xia, and Haibo Chen · 2024
Closest in time.
A survey on large language model (llm) security and privacy: The good, the bad, and the ugly
Yifan Yao, Jinhao Duan, Kaidi Xu, Yuanfang Cai, Zhibo Sun, and Yue Zhang · 2024
Closest in time.
Minicpm-v: A gpt-4v level mllm on your phone
Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, et al · 2024
Closest in time.
Llm as a system service on mobile devices
Wangsong Yin, Mengwei Xu, Yuanchun Li, and Xuanzhe Liu · 2024
Closest in time.
Hao Zhou, Chengming Hu, Ye Yuan, Yufei Cui, Yili Jin, Can Chen, Haolun Wu, Dun Yuan, Li Jiang, Di Wu, et al · 2024
Closest in time.
liburing
Jens Axboe · 2025
Closest in time.
Maxembed: Maximizing ssd bandwidth utilization for huge embedding models serving
Ruwen Fan, Minhui Xie, Haodi Jiang, and Youyou Lu · 2025
Closest in time.
Bitsandbytes
BitsandBytes Foundation · 2025
Closest in time.
Termux app
Termux · 2025
Closest in time.
Lemo: Enabling less token involvement for more context fine-tuning, 2025
Tuowei Wang, Xingyu Chen, Kun Li, Ting Cao, Ju Ren, and Yaoxue Zhang · 2025
Closest in time.