Fetching the paper…
Reading the bibliography…
On-device inference for Large Language Models (LLMs), driven by increasing privacy concerns and advancements of mobile-sized models, has gained significant interest.
Generating simd vectorized permutations
Franz Franchetti and Markus Püschel · 2008
Earlier work this paper cites.
Traveling salesman problem
Karla L Hoffman, Manfred Padberg, Giovanni Rinaldi, et al · 2013
Earlier work this paper cites.
Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines
Jonathan Ragan-Kelley, Connelly Barnes, Andrew Adams, Sylvain Paris, Frédo Durand, and Saman Amarasinghe · 2013
Earlier work this paper cites.
Dsp. ear: Leveraging co-processor support for continuous audio sensing on smartphones
Petko Georgiev, Nicholas D Lane, Kiran K Rachuri, and Cecilia Mascolo · 2014
Earlier work this paper cites.
Deepx: A software accelerator for low-power deep learning inference on mobile devices
Nicholas D Lane, Sourav Bhattacharya, Petko Georgiev, Claudio Forlivesi, Lei Jiao, Lorena Qendro, and Fahim Kawsar · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernandez · 2016
Earlier work this paper cites.
Deepmon: Mobile gpu-based deep learning framework for continuous vision applications
Loc N Huynh, Youngki Lee, and Rajesh Krishna Balan · 2017
Earlier work this paper cites.
Tensorflow-serving: Flexible, high-performance ml serving
Christopher Olston, Noah Fiedel, Kiril Gorovoy, Jeremiah Harmsen, Li Lao, Fangwei Li, Vinu Rajashekhar, Sukriti Ramesh, and Jordan Soyke · 2017
Earlier work this paper cites.
Extending halide to improve software development for imaging dsps
Sander Vocke, Henk Corporaal, Roel Jordans, Rosilde Corvino, and Rick Nas · 2017
Earlier work this paper cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Earlier work this paper cites.
Deepcache: Principled cache for mobile deep vision
Mengwei Xu, Mengze Zhu, Yunxin Liu, Felix Xiaozhu Lin, and Xuanzhe Liu · 2018
Earlier work this paper cites.
Mosaic: Heterogeneity-, communication-, and constraint-aware model slicing and execution for accurate and efficient inference
Myeonggyun Han, Jihoon Hyun, Seongbeom Park, Jinsu Park, and Woongki Baek · 2019
Earlier work this paper cites.
μ \mu layer: Low latency on-device inference using cooperative single-layer acceleration and processor-friendly quantization
Youngsok Kim, Joonsung Kim, Dongju Chae, Daehyun Kim, and Jangwoo Kim · 2019
Earlier work this paper cites.
Mobisr: Efficient on-device super-resolution through heterogeneous mobile processors
Royson Lee, Stylianos I Venieris, Lukasz Dudziak, Sourav Bhattacharya, and Nicholas D Lane · 2019
Earlier work this paper cites.
Deploy machine learning models on mobile and iot devices, 2019
TensorFlow Lite · 2019
Earlier work this paper cites.
Triton: an intermediate language and compiler for tiled neural network computations
Philippe Tillet, Hsiang-Tsung Kung, and David Cox · 2019
Earlier work this paper cites.
A first look at deep learning apps on smartphones
Mengwei Xu, Jiawei Liu, Yuanqiang Liu, Felix Xiaozhu Lin, Yunxin Liu, and Xuanzhe Liu · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Earlier work this paper cites.
Mnn: A universal and efficient inference engine
Xiaotang Jiang, Huan Wang, Yiliu Chen, Ziqi Wu, Lichuan Wang, Bin Zou, Yafeng Yang, Zongyang Cui, Yu Cai, Tianhang Yu, et al · 2020
Earlier work this paper cites.
Patdnn: Achieving real-time dnn execution on mobile devices with pattern-based weight pruning
Wei Niu, Xiaolong Ma, Sheng Lin, Shihao Wang, Xuehai Qian, Xue Lin, Yanzhi Wang, and Bin Ren · 2020
Earlier work this paper cites.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou · 2020
Earlier work this paper cites.
Heimdall: mobile gpu coordination platform for augmented reality applications
Juheon Yi and Youngki Lee · 2020
Earlier work this paper cites.
Mobipose: Real-time multi-person pose estimation on mobile devices
Jinrui Zhang, Deyu Zhang, Xiaohui Xu, Fucheng Jia, Yunxin Liu, Xuanzhe Liu, Ju Ren, and Yaoxue Zhang · 2020
Earlier work this paper cites.
https://gdpr-info.eu/ , 2021
General data protection regulation · 2021
Earlier work this paper cites.
Cocopie: Enabling real-time ai on off-the-shelf mobile devices via compression-compilation co-design
Hui Guan, Shaoshan Liu, Xiaolong Ma, Wei Niu, Bin Ren, Xipeng Shen, Yanzhi Wang, and Pu Zhao · 2021
Earlier work this paper cites.
Accelerating on-device learning with layer-wise processor selection method on unified memory
Donghee Ha, Mooseop Kim, KyeongDeok Moon, and Chi Yoon Jeong · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Dynamicvit: Efficient vision transformers with dynamic token sparsification
Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh · 2021
Earlier work this paper cites.
Spatten: Efficient sparse attention architecture with cascade token and head pruning
Hanrui Wang, Zhekai Zhang, and Song Han · 2021
Earlier work this paper cites.
Energy-efficient resource management for federated edge learning with cpu-gpu heterogeneous computing
Qunsong Zeng, Yuqing Du, Kaibin Huang, and Kin K Leung · 2021
Earlier work this paper cites.
Token merging: Your vit but faster
Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer, and Judy Hoffman · 2022
Earlier work this paper cites.
Enable deep learning on mobile devices: Methods, systems, and applications
Han Cai, Ji Lin, Yujun Lin, Zhijian Liu, Haotian Tang, Hanrui Wang, Ligeng Zhu, and Song Han · 2022
Earlier work this paper cites.
Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer · 2022
Earlier work this paper cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh · 2022
Earlier work this paper cites.
Learned token pruning for transformers
Sehoon Kim, Sheng Shen, David Thorsley, Amir Gholami, Woosuk Kwon, Joseph Hassoun, and Kurt Keutzer · 2022
Cited alongside, same era.
Gcd 2: A globally optimizing compiler for mapping dnns to mobile dsps
Wei Niu, Jiexiong Guan, Xipeng Shen, Yanzhi Wang, Gagan Agrawal, and Bin Ren · 2022
Cited alongside, same era.
Mandheling: Mixed-precision on-device dnn training with dsp offloading
Daliang Xu, Mengwei Xu, Qipeng Wang, Shangguang Wang, Yun Ma, Kang Huang, Gang Huang, Xin Jin, and Xuanzhe Liu · 2022
Cited alongside, same era.
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Zhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu, Conglong Li, and Yuxiong He · 2022
Cited alongside, same era.
Speechmoe2: Mixture-of-experts model with improved routing
Zhao You, Shulin Feng, Dan Su, and Dong Yu · 2022
Cited alongside, same era.
https://developer.android.com/ai/aicore
Empowering llm to use smartphone for intelligent task automation
Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby Jia-Jun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu · 2023
Later among the works it cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han · 2023
Later among the works it cites.
Niagara: Scheduling dnn inference services on heterogeneous edge processors
Daliang Xu, Qing Li, Mengwei Xu, Kang Huang, Gang Huang, Shangguang Wang, Xin Jin, Yun Ma, and Xuanzhe Liu · 2023
Later among the works it cites.
Llmcad: Fast and scalable on-device large language model inference
Daliang Xu, Wangsong Yin, Xin Jin, Ying Zhang, Shiyun Wei, Mengwei Xu, and Xuanzhe Liu · 2023
Later among the works it cites.
Predictive pipelined decoding: A compute-latency trade-off for exact llm decoding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
AI Core · 2023
Cited alongside, same era.
https://www.apple.com/apple-intelligence/
Apple Intelligence · 2023
Cited alongside, same era.
https://github.com/NVIDIA/FasterTransformer
Gboard - the Google Keyboard - Apps on Google Play — play.google.com · 2023
Cited alongside, same era.
https://www.sellcell.com/blog/how-often-do-people-upgrade-their-phone-2023-statistics , 2023
AMD Strix Point (Ryzen 300) · 2023
Cited alongside, same era.
https://www.hisilicon.com/en/products/Kirin/Kirin-flagship-chips/Kirin-9000 , 2023
Ascend NPU · 2023
Cited alongside, same era.
https://coral.ai/docs/edgetpu/inference/#general-purpose-operating-systems , 2023
Edge TPU API · 2023
Cited alongside, same era.
https://huggingface.co/google/gemma-2b , 2023
Gemma-2B · 2023
Cited alongside, same era.
Seongjun Yang, Gibbeum Lee, Jaewoong Cho, Dimitris Papailiopoulos, and Kangwook Lee · 2023
Later among the works it cites.
Edgemoe: Fast on-device inference of moe-based large language models
Rongjie Yi, Liwei Guo, Shiyun Wei, Ao Zhou, Shangguang Wang, and Mengwei Xu · 2023
Later among the works it cites.
https://ai-benchmark.com/ranking_detailed.html , 2024
Ai benchmark · 2024
Closest in time.
https://cloud.google.com/edge-tpu , 2024
Edgetpu · 2024
Closest in time.
https://hix.ai/ai-email-writer-email-generator , 2024
Gpt-based email writer · 2024
Closest in time.
https://huggingface.co/ , 2024
Hugging face · 2024
Closest in time.
https://github.com/LlamaTouch/LlamaTouch , 2024
Llamatouch · 2024
Closest in time.
https://github.com/UbiquitousLearning/mllm , 2024
Mllm · 2024
Closest in time.
https://paperswithcode.com/sota/multi-task-language-understanding-on-mmlu , 2024
Mmlu leader board · 2024
Closest in time.
https://www.qualcomm.com/developer/software/qualcomm-ai-engine-direct-sdk , 2024
Qnn · 2024
Closest in time.
Phi-3 technical report: A highly capable language model locally on your phone
Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl, et al · 2024
Closest in time.
Medusa: Simple llm inference acceleration framework with multiple decoding heads
Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D Lee, Deming Chen, and Tri Dao · 2024
Closest in time.
Break the sequential dependency of llm inference using lookahead decoding
Yichao Fu, Peter Bailis, Ion Stoica, and Hao Zhang · 2024
Closest in time.
Pantheon: Preemptible multi-dnn inference on mobile edge gpus
Lixiang Han, Zimu Zhou, and Zhenjiang Li · 2024
Closest in time.
Personal llm agents: Insights and survey about the capability, efficiency and security
Yuanchun Li, Hao Wen, Weijun Wang, Xiangyu Li, Yizhen Yuan, Guohong Liu, Jiacheng Liu, Wenxing Xu, Xiang Wang, Yi Sun, et al · 2024
Closest in time.
Small language models: Survey, measurements, and insights
Zhenyan Lu, Xiang Li, Dongqi Cai, Rongjie Yi, Fangming Liu, Xiwen Zhang, Nicholas D Lane, and Mengwei Xu · 2024
Closest in time.
Sod2: Statically optimizing dynamic deep neural network
Wei Niu, Gagan Agrawal, and Bin Ren · 2024
Closest in time.
Smartmem: Layout transformation elimination and adaptation for efficient dnn execution on mobile
Wei Niu, Md Musfiqur Rahman Sanim, Zhihao Shu, Jiexiong Guan, Xipeng Shen, Miao Yin, Gagan Agrawal, and Bin Ren · 2024
Closest in time.
Automatic generation of vectorizing compilers for customizable digital signal processors
Samuel Thomas and James Bornholt · 2024
Closest in time.
Autodroid: Llm-powered task automation in android
Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby Jia-Jun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu · 2024
Closest in time.
Loongserve: Efficiently serving long-context large language models with elastic sequence parallelism
Bingyang Wu, Shengyu Liu, Yinmin Zhong, Peng Sun, Xuanzhe Liu, and Xin Jin · 2024
Closest in time.
Wip: Efficient llm prefilling with mobile npu
Daliang Xu, Hao Zhang, Liming Yang, Ruiqi Liu, Mengwei Xu, and Xuanzhe Liu · 2024
Closest in time.
{ \{ FwdLLM } \} : Efficient federated finetuning of large language models with perturbed inferences
Mengwei Xu, Dongqi Cai, Yaozong Wu, Xiang Li, and Shangguang Wang · 2024
Closest in time.
A survey of resource-efficient llm and multimodal foundation models
Mengwei Xu, Wangsong Yin, Dongqi Cai, Rongjie Yi, Daliang Xu, Qipeng Wang, Bingyang Wu, Yihao Zhao, Chen Yang, Shihe Wang, et al · 2024
Closest in time.
Muxflow: Efficient gpu sharing in production-level clusters with more than 10,000 gpus
Liu Xuanzhe, Yihao Zhao, Shufan Liu, Xiang Li, Xin Zhu, Yibo Liu, and Xin Jin · 2024
Closest in time.
Powerinfer-2: Fast large language model inference on a smartphone
Zhenliang Xue, Yixin Song, Zeyu Mi, Le Chen, Yubin Xia, and Haibo Chen · 2024
Closest in time.
Llm as a system service on mobile devices
Wangsong Yin, Mengwei Xu, Yuanchun Li, and Xuanzhe Liu · 2024
Closest in time.
Elms: Elasticized large language models on mobile devices
Wangsong Yin, Rongjie Yi, Daliang Xu, Gang Huang, Mengwei Xu, and Xuanzhe Liu · 2024
Closest in time.
Mobile foundation model as firmware
Jinliang Yuan, Chen Yang, Dongqi Cai, Shihe Wang, Xin Yuan, Zeling Zhang, Xiang Li, Dingge Zhang, Hanzi Mei, Xianqing Jia, et al · 2024
Closest in time.
Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang · 2024
Closest in time.