Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) exhibit remarkable human-like predictive capabilities.
Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics
Chin-Yew Lin and Franz Josef Och · 2004
Earlier work this paper cites.
Adaptive Computation Time for Recurrent Neural Networks
Alex Graves · 2016
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence RNNs and beyond
Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, Çağlar Gu̇lçehre, and Bing Xiang · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Branchynet: Fast inference via early exiting from deep neural networks
Surat Teerapittayanon, Bradley McDanel, and H.T. Kung · 2016
Earlier work this paper cites.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Shashi Narayan, Shay B. Cohen, and Mirella Lapata · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training, 2018
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
Shallow-deep networks: Understanding and mitigating network overthinking
Yigitcan Kaya, Sanghyun Hong, and Tudor Dumitras · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, and et al · 2020
Earlier work this paper cites.
Depth-adaptive transformer
Maha Elbayad, Jiatao Gu, Edouard Grave, and Michael Auli · 2020
Earlier work this paper cites.
Dynabert: Dynamic bert with adaptive width and depth
Lu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang, Xiao Chen, and Qun Liu · 2020
Earlier work this paper cites.
Fastbert: a self-distilling bert with adaptive inference time
Weijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao, Haotang Deng, and Qi Ju · 2020
Earlier work this paper cites.
The right tool for the job: Matching model and instance complexities
Roy Schwartz, Gabriel Stanovsky, Swabha Swayamdipta, Jesse Dodge, and Noah A. Smith · 2020
Earlier work this paper cites.
DeeBERT: Dynamic early exiting for accelerating BERT inference
Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin · 2020
Earlier work this paper cites.
Bert loses patience: Fast and robust inference with early exit
Wangchunshu Zhou, Canwen Xu, Tao Ge, Julian McAuley, Ke Xu, and Furu Wei · 2020
Earlier work this paper cites.
Adaptive inference through early-exit networks: Design, challenges and directions
Stefanos Laskaridis, Alexandros Kouris, and Nicholas D. Lane · 2021
Earlier work this paper cites.
Accelerating BERT inference for sequence labeling via early-exit
Xiaonan Li, Yunfan Shao, Tianxiang Sun, Hang Yan, Xipeng Qiu, and Xuanjing Huang · 2021
Earlier work this paper cites.
Consistent accelerated inference via confident adaptive transformers
Tal Schuster, Adam Fisch, Tommi Jaakkola, and Regina Barzilay · 2021
Earlier work this paper cites.
Parallel detection for efficient video analytics at the edge
Yanzhao Wu, Ling Liu, and Ramana Kompella · 2021
Earlier work this paper cites.
Cloud-edge inference under communication constraints: Data quantization and early exit
Yu Gao, Wei Wang, Dezhi Wang, Huiqiong Wang, and Zhaoyang Zhang · 2022
Earlier work this paper cites.
TruthfulQA: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, and et al · 2022
Cited alongside, same era.
A comparative measurement study of deep learning as a service framework
Yanzhao Wu, Ling Liu, Calton Pu, Wenqi Cao, Semih Sahin, Wenqi Wei, and Qi Zhang · 2022
Cited alongside, same era.
Fast and robust early-exiting framework for autoregressive language models with synchronized parallel decoding
Sangmin Bae, Jongwoo Ko, Hwanjun Song, and Se-Young Yun · 2023
Cited alongside, same era.
SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference
Luciano Del Corro, Allie Del Giorno, Sahaj Agarwal, Bin Yu, Ahmed Awadallah, and Subhabrata Mukherjee · 2023
Cited alongside, same era.
Gemini: A Family of Highly Capable Multimodal Models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, and et al · 2023
Ee-llm: Large-scale training and inference of early-exit large language models with 3d parallelism
Yanxi Chen, Xuchen Pan, Yaliang Li, Bolin Ding, and Jingren Zhou · 2024
Closest in time.
Security and Privacy Challenges of Large Language Models: A Survey
Badhan Chandra Das, M. Hadi Amini, and Yanzhao Wu · 2024
Closest in time.
Enhancing on-device llm inference with historical cloud-based llm interactions
Yucheng Ding, Chaoyue Niu, Fan Wu, Shaojie Tang, Chengfei Lyu, and Guihai Chen · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, and et al · 2024
Closest in time.
Confidential Prompting: Protecting User Prompts from Cloud LLM Providers
In Gim, Caihua Li, and Lin Zhong · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Smartbert: a promotion of dynamic early exiting mechanism for accelerating bert inference
Boren Hu, Yun Zhu, Jiacheng Li, and Siliang Tang · 2023
Cited alongside, same era.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, and et al · 2023
Cited alongside, same era.
Rethinking learning rate tuning in the era of large language models
Hongpeng Jin, Wenqi Wei, Xuyu Wang, Wenbin Zhang, and Yanzhao Wu · 2023
Cited alongside, same era.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, and et al · 2023
Cited alongside, same era.
Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Hongyi Jin, Tianqi Chen, and Zhihao Jia · 2023
Cited alongside, same era.
A comprehensive overview of large language models
Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad Saqib, Saeed Anwar, Muhammad Usman, Nick Barnes, and Ajmal Mian · 2023
Cited alongside, same era.
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, and et al · 2023
Cited alongside, same era.
Closest in time.
Hybrid slm and llm for edge-cloud collaborative inference
Zixu Hao, Huiqiang Jiang, Shiqi Jiang, Ju Ren, and Ting Cao · 2024
Closest in time.
Billm: Pushing the limit of post-training quantization for llms
Wei Huang, Yangdong Liu, Haotong Qin, Ying Li, Shiming Zhang, Xianglong Liu, Michele Magno, and Xiaojuan Qi · 2024
Closest in time.
Adaptive deep neural network inference optimization with eenet
Fatih Ilhan, Ka-Ho Chow, Sihao Hu, Tiansheng Huang, Selim Tekin, Wenqi Wei, Yanzhao Wu, Myungjin Lee, Ramana Kompella, Hugo Latapie, Gaowen Liu, and Ling Liu · 2024
Closest in time.
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, and et al · 2024
Closest in time.
Efficiently distilling LLMs for edge applications
Achintya Kundu, Yu Chin Fabian Lim, Aaron Chew, Laura Wynter, Penny Chong, and Rhui Lee · 2024
Closest in time.
Awq: Activation-aware weight quantization for on-device llm compression and acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han · 2024
Closest in time.
The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
Shuming Ma, Hongyu Wang, Lingxiao Ma, Lei Wang, Wenhui Wang, Shaohan Huang, Li Dong, Ruiping Wang, Jilong Xue, and Furu Wei · 2024
Closest in time.
Chatrtx: Bringing generative ai to consumers with nvidia ai on RTX
NVIDIA · 2024
Closest in time.
Ollama: Open large language model api
Ollama · 2024
Closest in time.
Ee-tuning: An economical yet scalable solution for tuning early-exit large language models, 2024
Xuchen Pan, Yanxi Chen, Yaliang Li, Bolin Ding, and Jingren Zhou · 2024
Closest in time.
Splitwise: Efficient Generative LLM Inference Using Phase Splitting
Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Inigo Goiri, Saeed Maleki, and Ricardo Bianchini · 2024
Closest in time.
Mobile Edge Intelligence for Large Language Models: A Contemporary Survey
Guanqiao Qu, Qiyuan Chen, Wei Wei, Zheng Lin, Xianhao Chen, and Kaibin Huang · 2024
Closest in time.
PAPILLON: PrivAcy Preservation from Internet-based and Local Language MOdel ENsembles
Li Siyan, Vethavikashini Chithrra Raghuram, Omar Khattab, Julia Hirschberg, and Zhou Yu · 2024
Closest in time.
Investigating acceleration of LLaMA inference by enabling intermediate layer decoding via instruction tuning with ‘LITE’
Neeraj Varshney, Agneet Chatterjee, Mihir Parmar, and Chitta Baral · 2024
Closest in time.
A Survey of Resource-efficient LLM and Multimodal Foundation Models
Mengwei Xu, Wangsong Yin, Dongqi Cai, Rongjie Yi, Daliang Xu, Qipeng Wang, Bingyang Wu, Yihao Zhao, Chen Yang, Shihe Wang, Qiyang Zhang, Zhenyan Lu, Li Zhang, Shangguang Wang, Yuanchun Li, Yunxin Liu, Xin Jin, and Xuanzhe Liu · 2024
Closest in time.
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
Mingjin Zhang, Jiannong Cao, Xiaoming Shen, and Zeyang Cui · 2024
Closest in time.
A Survey on Efficient Inference for Large Language Models
Zixuan Zhou, Xuefei Ning, Ke Hong, Tianyu Fu, Jiaming Xu, Shiyao Li, Yuming Lou, Luning Wang, Zhihang Yuan, Xiuhong Li, Shengen Yan, Guohao Dai, Xiao-Ping Zhang, Yuhan Dong, and Yu Wang · 2024
Closest in time.