Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have revolutionized natural language processing with their exceptional understanding, synthesizing, and reasoning capabilities.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Demo: Video Analytics - Killer App for Edge Computing. In ACM MobiSys
Ganesh Ananthanarayanan, Victor Bahl, Landon Cox, Alex Crown, Shadi Nogbahi, and Yuanchao Shu. 2019 · 2019
Earlier work this paper cites.
Unified language model pre-training for natural language understanding and generation
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019 · 2019
Earlier work this paper cites.
Networked Cameras Are the New Big Data Clusters. In ACM Workshop on Hot Topics in Video Analytics and Intelligent Edges
Junchen Jiang, Yuhao Zhou, Ganesh Ananthanarayanan, Yuanchao Shu, and Andrew A. Chien. 2019 · 2019
Earlier work this paper cites.
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 2019
Earlier work this paper cites.
Collaborative Learning between Cloud and End Devices: An Empirical Study on Location Prediction. In ACM/IEEE SEC
Yan Lu, Yuanchao Shu, Xu Tan, Yunxin Liu, Mengyu Zhou, Qi Chen, and Dan Pei. 2019 · 2019
Earlier work this paper cites.
Quantifying Attention Flow in Transformers
Samira Abnar and Willem Zuidema. 2020 · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Power-bert: Accelerating bert inference via progressive word-vector elimination
Saurabh Goyal, Anamitra Roy Choudhury, Saurabh Raje, Venkatesan Chakaravarthy, Yogish Sabharwal, and Ashish Verma. 2020 · 2020
Earlier work this paper cites.
Spatula: Efficient cross-camera video analytics on large camera networks
Samvit Jain, Xun Zhang, Yuhao Zhou, Ganesh Ananthanarayanan, Junchen Jiang, Yuanchao Shu, Paramvir Bahl, and Joseph Gonzalez. 2020 · 2020
Earlier work this paper cites.
TinyBERT: Distilling BERT for Natural Language Understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020 · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2020
Earlier work this paper cites.
Efficient Transformer-based Large Scale Language Representations using Hardware-friendly Block Structured Pruning
Bingbing Li, Zhenglun Kong, Tianyun Zhang, Ji Li, Zhengang Li, Hang Liu, and Caiwen Ding. 2020 · 2020
Earlier work this paper cites.
FastBERT: A Self-distilling BERT with Adaptive Inference Time
Weijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao, Haotang Deng, and Qi Ju. 2020 · 2020
Earlier work this paper cites.
Movement Pruning: Adaptive Sparsity by Fine-Tuning
Victor Sanh, Thomas Wolf, and Alexander Rush. 2020 · 2020
Earlier work this paper cites.
Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT
Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer. 2020 · 2020
Earlier work this paper cites.
MobileBERT: A Compact Task-Agnostic BERT for Resource-Limited Devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020 · 2020
Earlier work this paper cites.
Algorithm-hardware co-design of adaptive floating-point encodings for resilient deep learning inference
Thierry Tambe, En-Yu Yang, Zishen Wan, Yuntian Deng, Vijay Janapa Reddi, Alexander Rush, David Brooks, and Gu-Yeon Wei. 2020 · 2020
Earlier work this paper cites.
MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020b · 2020
Earlier work this paper cites.
Convergence of Edge Computing and Deep Learning: A Comprehensive Survey
Xiaofei Wang, Yiwen Han, Victor C. M. Leung, Dusit Niyato, Xueqiang Yan, and Xu Chen. 2020a · 2020
Earlier work this paper cites.
Gobo: Quantizing attention-based nlp models for low latency and energy efficient inference
Ali Hadi Zadeh, Isak Edo, Omar Mohamed Awad, and Andreas Moshovos. 2020 · 2020
Earlier work this paper cites.
TernaryBERT: Distillation-aware Ultra-low Bit BERT
Wei Zhang, Lu Hou, Yichun Yin, Lifeng Shang, Xiao Chen, Xin Jiang, and Qun Liu. 2020 · 2020
Earlier work this paper cites.
BERT Loses Patience: Fast and Robust Inference with Early Exit
Wangchunshu Zhou, Canwen Xu, Tao Ge, Julian McAuley, Ke Xu, and Furu Wei. 2020 · 2020
Earlier work this paper cites.
Understanding and Overcoming the Challenges of Efficient Transformer Quantization
Yelysei Bondarenko, Markus Nagel, and Tijmen Blankevoort. 2021 · 2021
Earlier work this paper cites.
DRONE: Data-aware Low-rank Compression for Large NLP Models
Patrick Chen, Hsiang-Fu Yu, Inderjit Dhillon, and Cho-Jui Hsieh. 2021 · 2021
Earlier work this paper cites.
Is Attention Better Than Matrix Decomposition?
Zhengyang Geng, Meng-Hao Guo, Hongxu Chen, Xia Li, Ke Wei, and Zhouchen Lin. 2021 · 2021
Earlier work this paper cites.
A Survey on Recent Approaches for Natural Language Processing in Low-Resource Scenarios
Michael A. Hedderich, Lukas Lange, Heike Adel, Jannik Strötgen, and Dietrich Klakow. 2021 · 2021
Earlier work this paper cites.
LoRA: Low-Rank Adaptation of Large Language Models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021 · 2021
Earlier work this paper cites.
On the Distribution, Sparsity, and Inference-time Quantization of Attention Values in Transformers
Tianchu Ji, Shraddhan Jain, Michael Ferdman, Peter Milder, H Andrew Schwartz, and Niranjan Balasubramanian. 2021 · 2021
Earlier work this paper cites.
Flexible high-resolution object detection on edge devices with tunable latency
Shiqi Jiang, Zhiqi Lin, Yuanchun Li, Yuanchao Shu, and Yunxin Liu. 2021 · 2021
Earlier work this paper cites.
Length-Adaptive Transformer: Train Once with Length Drop, Use Anytime with Search
Gyuwan Kim and Kyunghyun Cho. 2021 · 2021
Earlier work this paper cites.
Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture
Liqiang Lu, Yicheng Jin, Hangrui Bi, Zizhang Luo, Peng Li, Tao Wang, and Yun Liang. 2021 · 2021
Earlier work this paper cites.
Delight: Deep and light-weight transformer
Sachin Mehta, Marjan Ghazvininejad, Srinivasan Iyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2021 · 2021
Earlier work this paper cites.
DNNFusion: Accelerating Deep Neural Networks Execution with Advanced Operator Fusion
Wei Niu, Jiexiong Guan, Yanzhi Wang, Gagan Agrawal, and Bin Ren. 2021 · 2021
Earlier work this paper cites.
EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference
Thierry Tambe, Coleman Hooper, Lillian Pentecost, Tianyu Jia, En-Yu Yang, Marco Donato, Victor Sanh, Paul Whatmough, Alexander M. Rush, David Brooks, and Gu-Yeon Wei. 2021 · 2021
Earlier work this paper cites.
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
Hanrui Wang, Zhekai Zhang, and Song Han. 2021b · 2021
Earlier work this paper cites.
MiniLMv2: Multi-Head Self-Attention Relation Distillation for Compressing Pretrained Transformers
Wenhui Wang, Hangbo Bao, Shaohan Huang, Li Dong, and Furu Wei. 2021a · 2021
Earlier work this paper cites.
LeeBERT: Learned Early Exit for BERT with Cross-Level Optimization
Wei Zhu. 2021 · 2021
Earlier work this paper cites.
Ekya: Continuous learning of video analytics models on edge compute servers
Romil Bhardwaj, Zhengxu Xia, Ganesh Ananthanarayanan, Junchen Jiang, Yuanchao Shu, Nikolaos Karianakis, Kevin Hsieh, Paramvir Bahl, and Ion Stoica. 2022 · 2022
Earlier work this paper cites.
MobiVQA: Efficient On-Device Visual Question Answering
Qingqing Cao, Prerna Khanna, Nicholas D Lane, and Aruna Balasubramanian. 2022 · 2022
Earlier work this paper cites.
LLM. int8 () 8-bit matrix multiplication for transformers at scale
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022 · 2022
Earlier work this paper cites.
A survey of natural language generation
Chenhe Dong, Yinghui Li, Haifan Gong, Miaoxin Chen, Junxin Li, Ying Shen, and Min Yang. 2022 · 2022
Earlier work this paper cites.
Optimal Brain Compression: A Framework for Accurate Post-Training Quantization and Pruning
Elias Frantar and Dan Alistarh. 2022 · 2022
Earlier work this paper cites.
OPTQ: Accurate Quantization for Generative Pre-trained Transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2022 · 2022
Earlier work this paper cites.
Ant: Exploiting adaptive numerical data type for low-bit deep neural network quantization
Cong Guo, Chen Zhang, Jingwen Leng, Zihan Liu, Fan Yang, Yunxin Liu, Minyi Guo, and Yuhao Zhu. 2022 · 2022
Earlier work this paper cites.
Language Model Compression with Weighted Low-Rank Factorization
Yen-Chang Hsu, Ting Hua, Sungen Chang, Qian Lou, Yilin Shen, and Hongxia Jin. 2022 · 2022
Earlier work this paper cites.
Learned token pruning for transformers
Sehoon Kim, Sheng Shen, David Thorsley, Amir Gholami, Woosuk Kwon, Joseph Hassoun, and Kurt Keutzer. 2022 · 2022
Earlier work this paper cites.
An empirical survey on long document summarization: Datasets, models, and metrics
Huan Yee Koh, Jiaxin Ju, Ming Liu, and Shirui Pan. 2022 · 2022
Earlier work this paper cites.
Accelerating Inference for Pretrained Language Models by Unified Multi-Perspective Early Exiting
Jun Kong, Jin Wang, Liang-Chih Yu, and Xuejie Zhang. 2022 · 2022
Earlier work this paper cites.
The Optimal BERT Surgeon: Scalable and Accurate Second-Order Pruning for Large Language Models
Eldar Kurtic, Daniel Campos, Tuan Nguyen, Elias Frantar, Mark Kurtz, Benjamin Fineran, Michael Goin, and Dan Alistarh. 2022 · 2022
Earlier work this paper cites.
Recurrent Neural Networks for Edge Intelligence: A Survey
Varsha S. Lalapura, J. Amudha, and Hariramn Selvamuruga Satheesh. 2022 · 2022
Earlier work this paper cites.
Turbo: Opportunistic enhancement for edge video analytics
Yan Lu, Shiqi Jiang, Ting Cao, and Yuanchao Shu. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Efficient acceleration of deep learning inference on resource-constrained edge devices: A review
Md Maruf Hossain Shuvo, Syed Kamrul Islam, Jianlin Cheng, and Bashir I Morshed. 2022 · 2022
Earlier work this paper cites.
A 28nm 15.59 μ \mu J/token full-digital bitline-transpose CIM-based sparse transformer accelerator with pipeline/parallel reconfigurable modes
Fengbin Tu, Zihan Wu, Yiqi Wang, Ling Liang, Liu Liu, Yufei Ding, Leibo Liu, Shaojun Wei, Yuan Xie, and Shouyi Yin. 2022 · 2022
Earlier work this paper cites.
Structured Pruning Learns Compact and Accurate Models
Mengzhou Xia, Zexuan Zhong, and Danqi Chen. 2022 · 2022
Earlier work this paper cites.
ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers
Zhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu, Conglong Li, and Yuxiong He. 2022 · 2022
Earlier work this paper cites.
Mokey: Enabling narrow fixed-point inference for out-of-the-box floating-point transformer models
Ali Hadi Zadeh, Mostafa Mahmoud, Ameer Abdelhadi, and Andreas Moshovos. 2022 · 2022
Earlier work this paper cites.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Earlier work this paper cites.
Transpim: A memory-based acceleration via software-hardware co-design for transformer
Minxuan Zhou, Weihong Xu, Jaeyoung Kang, and Tajana Rosing. 2022 · 2022
Earlier work this paper cites.
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebron, and Sumit Sanghai. 2023 · 2023
Earlier work this paper cites.
Fast and Robust Early-Exiting Framework for Autoregressive Language Models with Synchronized Parallel Decoding
Sangmin Bae, Jongwoo Ko, Hwanjun Song, and Se-Young Yun. 2023 · 2023
Earlier work this paper cites.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al · 2023
Earlier work this paper cites.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al · 2023
Earlier work this paper cites.
Petals: Collaborative Inference and Fine-tuning of Large Models
Alexander Borzunov, Dmitry Baranchuk, Tim Dettmers, Maksim Riabinin, Younes Belkada, Artem Chumachenko, Pavel Samygin, and Colin Raffel. 2023a · 2023
Earlier work this paper cites.
Distributed inference and fine-tuning of large language models over the internet
Alexander Borzunov, Max Ryabinin, Artem Chumachenko, Dmitry Baranchuk, Tim Dettmers, Younes Belkada, Pavel Samygin, and Colin Raffel. 2023b · 2023
Earlier work this paper cites.
MCC-KD: Multi-CoT Consistent Knowledge Distillation
Hongzhan Chen, Siyue Wu, Xiaojun Quan, Rui Wang, Ming Yan, and Ji Zhang. 2023b · 2023
Earlier work this paper cites.
Confidant: Customizing Transformer-based LLMs via Collaborative Edge Training
Yuhao Chen, Yuxuan Yan, Qianqian Yang, Yuanchao Shu, Shibo He, and Jiming Chen. 2023c · 2023
Earlier work this paper cites.
DISCO: Distilling Counterfactuals with Large Language Models
Zeming Chen, Qiyue Gao, Antoine Bosselut, Ashish Sabharwal, and Kyle Richardson. 2023a · 2023
Earlier work this paper cites.
Adapting Language Models to Compress Contexts
Alexis Chevalier, Alexander Wettig, Anirudh Ajith, and Danqi Chen. 2023 · 2023
Earlier work this paper cites.
SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
Tim Dettmers, Ruslan A. Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan Alistarh. 2023 · 2023
Earlier work this paper cites.
TaskFusion: An Efficient Transfer Learning Architecture with Dual Delta Sparsity for Multi-Task Natural Language Processing
Zichen Fan, Qirui Zhang, Pierre Abillama, Sara Shoouri, Changwoo Lee, David Blaauw, Hun-Seok Kim, and Dennis Sylvester. 2023 · 2023
Earlier work this paper cites.
SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
Elias Frantar and Dan Alistarh. 2023 · 2023
Earlier work this paper cites.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao. 2023 · 2023
Earlier work this paper cites.
MiniLLM: Knowledge Distillation of Large Language Models
Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang. 2023 · 2023
Earlier work this paper cites.
Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, et al · 2023
Earlier work this paper cites.
OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization
Cong Guo, Jiaming Tang, Weiming Hu, Jingwen Leng, Chen Zhang, Fan Yang, Yunxin Liu, Minyi Guo, and Yuhao Zhu. 2023b · 2023
Earlier work this paper cites.
STI: Turbocharge NLP Inference at the Edge via Elastic Pipelining
Liwei Guo, Wonkyo Choe, and Felix Xiaozhu Lin. 2023a · 2023
Earlier work this paper cites.
Large Language Models Are Reasoning Teachers
Namgyu Ho, Laura Schmid, and Se-Young Yun. 2023 · 2023
Cited alongside, same era.
3d-llm: Injecting the 3d world into large language models
Yining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng, Yilun Du, Zhenfang Chen, and Chuang Gan. 2023 · 2023
Cited alongside, same era.
Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes
Cheng-Yu Hsieh, Chun-Liang Li, Chih-kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alex Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister. 2023 · 2023
Cited alongside, same era.
Inner Monologue: Embodied Reasoning through Planning with Language Models
Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, et al · 2023
Cited alongside, same era.
Winclip: Zero-/few-shot anomaly classification and segmentation
Jongheon Jeong, Yang Zou, Taewan Kim, Dongqing Zhang, Avinash Ravichandran, and Onkar Dabeer. 2023 · 2023
Cited alongside, same era.
Biomistral: A collection of open-source pretrained large language models for medical domains
Yanis Labrak, Adrien Bazoge, Emmanuel Morin, Pierre-Antoine Gourraud, Mickael Rouvier, and Richard Dufour. 2024 · 2024
Closest in time.
A survey on automatic generation of figurative language: From rule-based systems to large language models
Huiyuan Lai and Malvina Nissim. 2024 · 2024
Closest in time.
MELTing Point: Mobile Evaluation of Language Transformers
Stefanos Laskaridis, Kleomenis Kateveas, Lorenzo Minto, and Hamed Haddadi. 2024 · 2024
Closest in time.
OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models
Changhun Lee, Jungyu Jin, Taesu Kim, Hyungjun Kim, and Eunhyeok Park. 2024b · 2024
Closest in time.
An Autonomous Parallelization of Transformer Model Inference on Heterogeneous Edge Devices
Juhyeon Lee, Insung Bahk, Hoseung Kim, Sinjin Jeong, Suyeon Lee, and Donghyun Min. 2024a · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models
Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2023b · 2023
Cited alongside, same era.
Lion: Adversarial Distillation of Proprietary Large Language Models
Yuxin Jiang, Chunkit Chan, Mingyang Chen, and Wei Wang. 2023a · 2023
Cited alongside, same era.
RECL: Responsive Resource-Efficient continuous learning for video analytics
Mehrdad Khani, Ganesh Ananthanarayanan, Kevin Hsieh, Junchen Jiang, Ravi Netravali, Yuanchao Shu, Mohammad Alizadeh, and Victor Bahl. 2023 · 2023
Cited alongside, same era.
Token-scaled logit distillation for ternary weight generative language models
Minsoo Kim, Sihwa Lee, Janghwan Lee, Sukjin Hong, Du-Seong Chang, Wonyong Sung, and Jungwook Choi. 2023 · 2023
Cited alongside, same era.
Efficient Memory Management for Large Language Model Serving with PagedAttention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Cited alongside, same era.
Fast inference from transformers via speculative decoding
Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023 · 2023
Cited alongside, same era.
Symbolic Chain-of-Thought Distillation: Small Models Can Also “Think” Step-by-Step
Liunian Harold Li, Jack Hessel, Youngjae Yu, Xiang Ren, Kai-Wei Chang, and Yejin Choi. 2023b · 2023
Cited alongside, same era.
SheetCopilot: Bringing software productivity to the next level through large language models
Hongxin Li, Jingran Su, Yuntao Chen, Qing Li, and ZHAO-XIANG ZHANG. 2024a · 2024
Closest in time.
Personal llm agents: Insights and survey about the capability, efficiency and security
Yuanchun Li, Hao Wen, Weijun Wang, Xiangyu Li, Yizhen Yuan, Guohong Liu, Jiacheng Liu, Wenxing Xu, Xiang Wang, Yi Sun, et al · 2024
Closest in time.
Jamba: A hybrid transformer-mamba language model
Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov, Shai Shalev-Shwartz, et al · 2024
Closest in time.
AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. 2024 · 2024
Closest in time.
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al · 2024
Closest in time.
Mobilellm: Optimizing sub-billion parameter language models for on-device use cases
Zechun Liu, Changsheng Zhao, Forrest Iandola, Chen Lai, Yuandong Tian, Igor Fedorov, Yunyang Xiong, Ernie Chang, Yangyang Shi, Raghuraman Krishnamoorthi, et al · 2024
Closest in time.
A multimodal generative AI copilot for human pathology
Ming Y Lu, Bowen Chen, Drew FK Williamson, Richard J Chen, Melissa Zhao, Aaron K Chow, Kenji Ikemura, Ahrong Kim, Dimitra Pouli, Ankush Patel, et al · 2024
Closest in time.
HPipe: Large Language Model Pipeline Parallelism for Long Context on Heterogeneous Cost-effective Devices
Ruilong Ma, Xiang Yang, Jingyu Wang, Qi Qi, Haifeng Sun, Jing Wang, Zirui Zhuang, and Jianxin Liao. 2024a · 2024
Closest in time.
Coco-agent: A comprehensive cognitive mllm agent for smartphone gui automation
Xinbei Ma, Zhuosheng Zhang, and Hai Zhao. 2024b · 2024
Closest in time.
OpenELM: An Efficient Language Model Family with Open-source Training and Inference Framework
Sachin Mehta, Mohammad Hossein Sekhavat, Qingqing Cao, Maxwell Horton, Yanzi Jin, Chenfan Sun, Iman Mirzadeh, Mahyar Najibi, Dmitry Belenko, Peter Zatloukal, et al · 2024
Closest in time.
ONNX Runtime is a cross-platform inference and training machine-learning accelerator
Microsoft. 2018 · 2024
Closest in time.
Phi-2: The Surprising Power of Small Language Models
Sébastien Bubeck Mojan Javaheripi. 2023 · 2024
Closest in time.
SmartMem: Layout Transformation Elimination and Adaptation for Efficient DNN Execution on Mobile
Wei Niu, Md Musfiqur Rahman Sanim, Zhihao Shu, Jiexiong Guan, Xipeng Shen, Miao Yin, Gagan Agrawal, and Bin Ren. 2024 · 2024
Closest in time.
TensorRT-LLM
NVIDIA. 2023 · 2024
Closest in time.
Minitron-4B-Base
NVIDIA. 2024 · 2024
Closest in time.
Introducing ChatGPT
OpenAI. 2022 · 2024
Closest in time.
Propagating knowledge updates to lms through distillation
Shankar Padmanabhan, Yasumasa Onoe, Michael Zhang, Greg Durrett, and Eunsol Choi. 2024 · 2024
Closest in time.
VLP: Vision Language Planning for Autonomous Driving
Chenbin Pan, Burhaneddin Yaman, Tommaso Nesti, Abhirup Mallik, Alessandro G Allievi, Senem Velipasalar, and Liu Ren. 2024 · 2024
Closest in time.
Nemotron-4 15B Technical Report
Jupinder Parmar, Shrimai Prabhumoye, Joseph Jennings, Mostofa Patwary, Sandeep Subramanian, Dan Su, Chen Zhu, Deepak Narayanan, Aastha Jhunjhunwala, Ayush Dattagupta, et al · 2024
Closest in time.
Raspberry Pi 4 on sale now
Raspberry Pi. 2019 · 2024
Closest in time.
ExecuTorch
Pytorch. 2023 · 2024
Closest in time.
Interactive continual learning: Fast and slow thinking
Biqing Qi, Xinquan Chen, Junqi Gao, Dong Li, Jianxing Liu, Ligang Wu, and Bowen Zhou. 2024 · 2024
Closest in time.
Mobile Edge Intelligence for Large Language Models: A Contemporary Survey
Guanqiao Qu, Qiyuan Chen, Wei Wei, Zheng Lin, Xianhao Chen, and Kaibin Huang. 2024 · 2024
Closest in time.
Snapdragon 8 Series Mobile Platforms | Qualcomm
Qualcomm. 2024 · 2024
Closest in time.
Pixel 8 Pro — the first smartphone with AI built in — is now running Gemini Nano, plus more AI updates coming to the Pixel portfolio
Brian Rakowski. 2023 · 2024
Closest in time.
Observational Scaling Laws and the Predictability of Language Model Performance
Yangjun Ruan, Chris J Maddison, and Tatsunori Hashimoto. 2024 · 2024
Closest in time.
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
Xuan Shen, Peiyan Dong, Lei Lu, Zhenglun Kong, Zhengang Li, Ming Lin, Chao Wu, and Yanzhi Wang. 2024 · 2024
Closest in time.
Powerinfer: Fast large language model serving with a consumer-grade gpu
Yixin Song, Zeyu Mi, Haotong Xie, and Haibo Chen. 2024 · 2024
Closest in time.
RoFormer: Enhanced Transformer with Rotary Position Embedding
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2024 · 2024
Closest in time.
A Simple and Effective Pruning Approach for Large Language Models
Mingjie Sun, Zhuang Liu, Anna Bair, and J. Zico Kolter. 2024a · 2024
Closest in time.
Spectr: Fast speculative decoding via optimal transport
Ziteng Sun, Ananda Theertha Suresh, Jae Hun Ro, Ahmad Beirami, Himanshu Jain, and Felix Yu. 2024b · 2024
Closest in time.
Graphgpt: Graph instruction tuning for large language models
Jiabin Tang, Yuhao Yang, Wei Wei, Lei Shi, Lixin Su, Suqi Cheng, Dawei Yin, and Chao Huang. 2024 · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al · 2024
Closest in time.
Snapdragon 8 Gen 3 Mobile Platform
Qualcomm Team. 2024 · 2024
Closest in time.
Large-Scale Deterministic Networks: Architecture, Enabling Technologies, Case Study and Future Directions
Wenbin Tian, Chaojie Gu, Miao Guo, Shibo He, Jiawen Kang, Dusit Niyato, and Jiming Chen. 2024a · 2024
Closest in time.
Drivevlm: The convergence of autonomous driving and large vision-language models
Xiaoyu Tian, Junru Gu, Bailin Li, Yicheng Liu, Yang Wang, Zhiyong Zhao, Kun Zhan, Peng Jia, Xianpeng Lang, and Hang Zhao. 2024b · 2024
Closest in time.
Dialogue summarization with mixture of experts based on large language models
Yuanhe Tian, Fei Xia, and Yan Song. 2024c · 2024
Closest in time.
Quip#: Even better LLM quantization with hadamard incoherence and lattice codebooks
Albert Tseng, Jerry Chee, Qingyao Sun, Volodymyr Kuleshov, and Christopher De Sa. 2024 · 2024
Closest in time.
Knowledge fusion of large language models
Fanqi Wan, Xinting Huang, Deng Cai, Xiaojun Quan, Wei Bi, and Shuming Shi. 2024a · 2024
Closest in time.
Efficient Large Language Models: A Survey
Zhongwei Wan, Xin Wang, Che Liu, Samiul Alam, Yu Zheng, Jiachen Liu, Zhongnan Qu, Shen Yan, Yi Zhu, Quanlu Zhang, Mosharaf Chowdhury, and Mi Zhang. 2024b · 2024
Closest in time.
End-edge-cloud collaborative computing for deep learning: A comprehensive survey
Yingchao Wang, Chen Yang, Shulin Lan, Liehuang Zhu, and Yan Zhang. 2024 · 2024
Closest in time.
AutoDroid: LLM-powered Task Automation in Android
Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby Jia-Jun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu. 2024 · 2024
Closest in time.
LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale Instructions
Minghao Wu, Abdul Waheed, Chiyu Zhang, Muhammad Abdul-Mageed, and Alham Aji. 2024 · 2024
Closest in time.
Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
Mengzhou Xia, Tianyu Gao, Zhiyuan Zeng, and Danqi Chen. 2024 · 2024
Closest in time.
Language models meet world models: Embodied experiences enhance language models
Jiannan Xiang, Tianhua Tao, Yi Gu, Tianmin Shu, Zirui Wang, Zichao Yang, and Zhiting Hu. 2024 · 2024
Closest in time.
EdgeLLM: Fast On-device LLM Inference with Speculative Decoding
Daliang Xu, Wangsong Yin, Hao Zhang, Xin Jin, Ying Zhang, Shiyun Wei, Mengwei Xu, and Xuanzhe Liu. 2024d · 2024
Closest in time.
GesturePrint: Enabling user identification for mmWave-based gesture recognition systems
Lilin Xu, Keyi Wang, Chaojie Gu, Xiuzhen Guo, Shibo He, and Jiming Chen. 2024b · 2024
Closest in time.
BESA: Pruning Large Language Models with Blockwise Parameter-Efficient Sparsity Allocation
Peng Xu, Wenqi Shao, Mengzhao Chen, Shitao Tang, Kaipeng Zhang, Peng Gao, Fengwei An, Yu Qiao, and Ping Luo. 2024a · 2024
Closest in time.
Mental-llm: Leveraging large language models for mental health prediction via online text data
Xuhai Xu, Bingsheng Yao, Yuanzhe Dong, Saadia Gabriel, Hong Yu, James Hendler, Marzyeh Ghassemi, Anind K Dey, and Dakuo Wang. 2024c · 2024
Closest in time.
Yuxuan Yan, Qianqian Yang, Shunpu Tang, and Zhiguo Shi. 2024 · 2024
Closest in time.
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al · 2024
Closest in time.
Large language models for test-free fault localization
Aidan ZH Yang, Claire Le Goues, Ruben Martins, and Vincent Hellendoorn. 2024b · 2024
Closest in time.
MentaLLaMA: interpretable mental health analysis on social media with large language models
Kailai Yang, Tianlin Zhang, Ziyan Kuang, Qianqian Xie, Jimin Huang, and Sophia Ananiadou. 2024d · 2024
Closest in time.
MAF: Exploring Mobile Acoustic Field for Hand-to-Face Gesture Interactions
Yongjie Yang, Tao Chen, Yujing Huang, Xiuzhen Guo, and Longfei Shangguan. 2024a · 2024
Closest in time.
Mobile Foundation Model as Firmware
Jinliang Yuan, Chen Yang, Dongqi Cai, Shihe Wang, Xin Yuan, Zeling Zhang, Xiang Li, Dingge Zhang, Hanzi Mei, Xianqing Jia, et al · 2024
Closest in time.
ConsistentEE: A Consistent and Hardness-Guided Early Exiting Method for Accelerating Language Models Inference
Ziqian Zeng, Yihuai Hong, Hongliang Dai, Huiping Zhuang, and Cen Chen. 2024 · 2024
Closest in time.
LlamaTouch: A Faithful and Scalable Testbed for Mobile UI Automation Task Evaluation
Li Zhang, Shihe Wang, Xianqing Jia, Zhihan Zheng, Yunhe Yan, Longxi Gao, Yuanchun Li, and Mengwei Xu. 2024e · 2024
Closest in time.
LoRAPrune: Structured Pruning Meets Low-Rank Parameter-Efficient Fine-Tuning
Mingyang Zhang, Hao Chen, Chunhua Shen, Zhen Yang, Linlin Ou, Xinyi Yu, and Bohan Zhuang. 2024b · 2024
Closest in time.
Beyond the Cloud: Edge Inference for Generative Large Language Models in Wireless Networks
Xinyuan Zhang, Jiangtian Nie, Yudong Huang, Gaochang Xie, Zehui Xiong, Jiang Liu, Dusit Niyato, and Xuemin Sherman Shen. 2024d · 2024
Closest in time.
Plug-and-play: An efficient post-training pruning method for large language models
Yingtao Zhang, Haoli Bai, Haokun Lin, Jialin Zhao, Lu Hou, and Carlo Vittorio Cannistraci. 2024a · 2024
Closest in time.
Vulcan: Automatic Query Planning for Live ML Analytics
Yiwen Zhang, Xumiao Zhang, Ganesh Ananthanarayanan, Anand Iyer, Yuanchao Shu, Victor Bahl, Z Morley Mao, and Mosharaf Chowdhury. 2024f · 2024
Closest in time.
Q-Hitter: A Better Token Oracle for Efficient LLM Inference via Sparse-Quantized KV Cache
Zhenyu Zhang, Shiwei Liu, Runjin Chen, Bhavya Kailkhura, Beidi Chen, and Atlas Wang. 2024c · 2024
Closest in time.
LinguaLinked: Distributed Large Language Model Inference on Mobile Devices
Junchen Zhao, Yurun Song, Ian Harris, Sangeetha Abdu Jyothi, et al · 2024
Closest in time.
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang. 2024 · 2024
Closest in time.
AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection
Qihang Zhou, Guansong Pang, Yu Tian, Shibo He, and Jiming Chen. 2024 · 2024
Closest in time.
Do LLMs Understand Visual Anomalies? Uncovering LLM’s Capabilities in Zero-shot Anomaly Detection
Jiaqi Zhu, Shaofeng Cai, Fang Deng, Beng Chin Ooi, and Junran Wu. 2024a · 2024
Closest in time.
Towards an On-device Agent for Text Rewriting
Yun Zhu, Yinxiao Liu, Felix Stahlberg, Shankar Kumar, Yu-Hui Chen, Liangchen Luo, Lei Shu, Renjie Liu, Jindong Chen, and Lei Meng. 2024b · 2024
Closest in time.
Eros: Real-time Dense Mapping Made Easy on Mobile Devices. In ACM HotMobile
Yubin Dai, Bin Qian, Yangkun Liu, Yuxuan Yan, and Yuanchao Shu. 2025 · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
End-to-End Multitarget Flexible Job Shop Scheduling With Deep Reinforcement Learning
Rongkai Wang, Yiyang Jing, Chaojie Gu, Shibo He, and Jiming Chen. 2025a · 2025
Closest in time.
Fed-DFA: Federated Distillation for Heterogeneous Model Fusion through the Adversarial Lens
Zichen Wang, Feng Yan, Tianyi Wang, Cong Wang, Yuanchao Shu, Peng Cheng, and Jiming Chen. 2025b · 2025
Closest in time.