Fetching the paper…
Reading the bibliography…
With the rapid advancement of artificial intelligence technologies such as ChatGPT, AI agents, and video generation, contemporary mobile systems have begun integrating these AI capabilities on local devices to enhance privacy and reduce response latency.
Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines
Jonathan Ragan-Kelley, Connelly Barnes, Andrew Adams, Sylvain Paris, Frédo Durand, and Saman Amarasinghe · 2013
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2016
Earlier work this paper cites.
Theano: A python framework for fast computation of mathematical expressions
The Theano Development Team, Rami Al-Rfou, Guillaume Alain, Amjad Almahairi, Christof Angermueller, Dzmitry Bahdanau, Nicolas Ballas, Frédéric Bastien, Justin Bayer, Anatoly Belikov, et al · 2016
Earlier work this paper cites.
TVM: An automated End-to-End optimizing compiler for deep learning
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy · 2018
Earlier work this paper cites.
Optimizing CNN model inference on CPUs
Yizhi Liu, Yao Wang, Ruofei Yu, Mu Li, Vin Sharma, and Yida Wang · 2019
Earlier work this paper cites.
MArk: Exploiting cloud services for Cost-Effective, SLO-Aware machine learning inference serving
Chengliang Zhang, Minchen Yu, Wei Wang, and Feng Yan · 2019
Earlier work this paper cites.
Accpar: Tensor partitioning for heterogeneous deep learning accelerators
Linghao Song, Fan Chen, Youwei Zhuo, Xuehai Qian, Hai Li, and Yiran Chen · 2020
Earlier work this paper cites.
Ansor: Generating High-Performance tensor programs for deep learning
Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, Joseph E. Gonzalez, and Ion Stoica · 2020
Earlier work this paper cites.
Accelerating applications using edge tensor processing units
Kuan-Chieh Hsu and Hung-Wei Tseng · 2021
Earlier work this paper cites.
Heterogeneous dataflow accelerators for multi-dnn workloads
Hyoukjun Kwon, Liangzhen Lai, Michael Pellauer, Tushar Krishna, Yu-Hsin Chen, and Vikas Chandra · 2021
Earlier work this paper cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh · 2022
Earlier work this paper cites.
A unified programmable edge matrix processor for deep neural networks and matrix algebra
Biji George, Om Ji Omer, Ziaul Choudhury, Anoop V, and Sreenivas Subramoney · 2022
Earlier work this paper cites.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang · 2022
Earlier work this paper cites.
Whale: Efficient giant model training over heterogeneous { \{ GPUs } \}
Xianyan Jia, Le Jiang, Ang Wang, Wencong Xiao, Ziji Shi, Jie Zhang, Xinyuan Li, Langshi Chen, Yong Li, Zhen Zheng, et al · 2022
Earlier work this paper cites.
Gunho Park, Baeseong Park, Minsub Kim, Sungjae Lee, Jeonghoon Kim, Beomseok Kwon, Se Jung Kwon, Byeongwook Kim, Youngjoo Lee, and Dongsoo Lee · 2022
Earlier work this paper cites.
Orca: A distributed serving system for { \{ Transformer-Based } \} generative models
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun · 2022
Earlier work this paper cites.
ROLLER: Fast and efficient tensor compilation for deep learning
Hongyu Zhu, Ruofan Wu, Yijia Diao, Shanbin Ke, Haoyu Li, Chen Zhang, Jilong Xue, Lingxiao Ma, Yuqing Xia, Wei Cui, Fan Yang, Mao Yang, Lidong Zhou, Asaf Cidon, and Gennady Pekhimenko · 2022
Earlier work this paper cites.
Teq: Trainable equivalent transformation for quantization of llms
Wenhua Cheng, Yiyang Cai, Kaokao Lv, and Haihao Shen · 2023
Earlier work this paper cites.
Simultaneous and heterogenous multithreading
Kuan-Chieh Hsu and Hung-Wei Tseng · 2023
Earlier work this paper cites.
Automated backend allocation for multi-model, on-device ai inference
Venkatraman Iyer, Sungho Lee, Semun Lee, Juitem Joonwoo Kim, Hyunjun Kim, and Youngjae Shin · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica · 2023
Earlier work this paper cites.
AlpaServe: Statistical multiplexing with model parallelism for deep learning serving
Zhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu, Ying Sheng, Xin Jin, Yanping Huang, Zhifeng Chen, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Cited alongside, same era.
Timothy R McIntosh, Teo Susnjak, Tong Liu, Paul Watters, and Malka N Halgamuge · 2023
Cited alongside, same era.
Accelerated edge machine learning
Microsoft · 2023
Cited alongside, same era.
Adainf: Data drift adaptive scheduling for accurate and slo-guaranteed multiple-model inference serving at edge servers
Sudipta Saha Shubha and Haiying Shen · 2023
Cited alongside, same era.
System virtualization for neural processing units
Yuqi Xue, Yiqi Liu, and Jian Huang · 2023
Cited alongside, same era.
Scar: Scheduling multi-model ai workloads on heterogeneous multi-chiplet module accelerators
Mohanad Odema, Luke Chen, Hyoukjun Kwon, and Mohammad Abdullah Al Faruque · 2024
Later among the works it cites.
Gsm8k (grade school math 8k)
OpenAI · 2024
Later among the works it cites.
Introducing chatgpt
OpenAI · 2024
Later among the works it cites.
Splitwise: Efficient generative llm inference using phase splitting
Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini · 2024
Later among the works it cites.
Snapdragon 8Gen3 - Mobile Platform ignites endless possibilities
Qualcomm · 2024
Later among the works it cites.
Llama-v3.1-8b-chat on qualcomm 8 elite
Qualcomm · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V10: Hardware-assisted npu multi-tenancy for improved resource utilization and fairness
Yuqi Xue, Yiqi Liu, Lifeng Nai, and Jian Huang · 2023
Cited alongside, same era.
Taming Throughput-Latency tradeoff in LLM inference with Sarathi-Serve
Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav Gulavani, Alexey Tumanov, and Ramachandran Ramjee · 2024
Cited alongside, same era.
Claude - right-sized for any task, the claude family of models offers the best combination of speed and performance
Anthropic · 2024
Cited alongside, same era.
Longbench: A bilingual, multitask benchmark for long context understanding, 2024
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li · 2024
Cited alongside, same era.
ServerlessLLM: Low-Latency serverless inference for large language models
Yao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete, Dmitrii Ustiugov, Yuvraj Patel, and Luo Mai · 2024
Cited alongside, same era.
llama.cpp - LLM inference with minimal setup and state-of-the-art performance on a wide range of hardware
Ggerganov · 2024
Cited alongside, same era.
Chatglm: A family of large language models from glm-130b to glm-4 all tools, 2024
Team GLM, Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Diego Rojas, Guanyu Feng, Hanlin Zhao, Hanyu Lai, Hao Yu, Hongning Wang, Jiadai Sun, Jiajie Zhang, Jiale Cheng, Jiayi Gui, Jie Tang, Jing Zhang, Juanzi Li, Lei Zhao, Lindong Wu, Lucen Zhong, Mingdao Liu, Minlie Huang, Peng Zhang, Qinkai Zheng, Rui Lu, Shuaiqi Duan, Shudan Zhang, Shulin Cao, Shuxun Yang, Weng Lam Tam, Wenyi Zhao, Xiao Liu, Xiao Xia, Xiaohan Zhang, Xiaotao Gu, Xin Lv, Xinghan Liu, Xinyi Liu, Xinyue Yang, Xixuan Song, Xunkai Zhang, Yifan An, Yifan Xu, Yilin Niu, Yuantao Yang, Yueyan Li, Yushi Bai, Yuxiao Dong, Zehan Qi, Zhaoyu Wang, Zhen Yang, Zhengxiao Du, Zhenyu Hou, and Zihan Wang · 2024
Cited alongside, same era.
NCNN - High-performance neural network inference computing framework optimized for mobile platforms
Tencent · 2024
Later among the works it cites.
Junyang Wang, Haiyang Xu, Haitao Jia, Xi Zhang, Ming Yan, Weizhou Shen, Ji Zhang, Fei Huang, and Jitao Sang · 2024
Later among the works it cites.
Mobile-agent: Autonomous multi-modal mobile device agent with visual perception
Junyang Wang, Haiyang Xu, Jiabo Ye, Ming Yan, Weizhou Shen, Ji Zhang, Fei Huang, and Jitao Sang · 2024
Later among the works it cites.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin · 2024
Later among the works it cites.
Mnn-llm: A generic inference engine for fast large language model deployment on mobile devices
Zhaode Wang, Jingbang Yang, Xinyu Qian, Shiwen Xing, Xiaotang Jiang, Chengfei Lv, and Shengyu Zhang · 2024
Later among the works it cites.
Fast on-device llm inference with npus, 2024
Daliang Xu, Hao Zhang, Liming Yang, Ruiqi Liu, Gang Huang, Mengwei Xu, and Xuanzhe Liu · 2024
Later among the works it cites.
Hardware-assisted virtualization of neural processing units for cloud platforms
Yuqi Xue, Yiqi Liu, Lifeng Nai, and Jian Huang · 2024
Later among the works it cites.
Powerinfer-2: Fast large language model inference on a smartphone, 2024
Zhenliang Xue, Yixin Song, Zeyu Mi, Xinrui Zheng, Yubin Xia, and Haibo Chen · 2024
Later among the works it cites.
Cogvideox: Text-to-video diffusion models with an expert transformer
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al · 2024
Later among the works it cites.
Pan Zhang, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Rui Qian, Lin Chen, Qipeng Guo, Haodong Duan, Bin Wang, Linke Ouyang, et al · 2024
Later among the works it cites.
DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang · 2024
Later among the works it cites.
MonoNN: Enabling a new monolithic optimization space for neural network inference tasks on modern GPU-Centric architectures
Donglin Zhuang, Zhen Zheng, Haojun Xia, Xiafei Qiu, Junjie Bai, Wei Lin, and Shuaiwen Leon Song · 2024
Later among the works it cites.
Qualcomm adreno gpu, game-changing speed and efficiency
Qualcomm · 2025
Closest in time.
Qualcomm hexagon npu, powering the generative ai revolution
Qualcomm · 2025
Closest in time.
Ferret-ui: Grounded mobile ui understanding with multimodal llms
Keen You, Haotian Zhang, Eldon Schoop, Floris Weers, Amanda Swearngin, Jeffrey Nichols, Yinfei Yang, and Zhe Gan · 2025
Closest in time.