Fetching the paper…
Reading the bibliography…
Large language models (LLMs) such as GPTs and Mixtral-8x7B have revolutionized machine intelligence due to their exceptional abilities in generic ML tasks.
ªadaptive mixtures of local experts, º neural computation, vol. 3
RA Jacobs, MI Jordan, SJ Nowlan, and GE Hinton · 1991
Earlier work this paper cites.
Recall-oriented understudy for gisting evaluation (rouge)
C Lin · 2005
Earlier work this paper cites.
Hierarchical models in the brain
Karl Friston · 2008
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally · 2015
Earlier work this paper cites.
Sparsification and separation of deep learning layers for constrained resource inference on wearables
Sourav Bhattacharya and Nicholas D Lane · 2016
Earlier work this paper cites.
Deepx: A software accelerator for low-power deep learning inference on mobile devices
Nicholas D Lane, Sourav Bhattacharya, Petko Georgiev, Claudio Forlivesi, Lei Jiao, Lorena Qendro, and Fahim Kawsar · 2016
Earlier work this paper cites.
Mobirnn: Efficient recurrent neural network execution on mobile gpu
Qingqing Cao, Niranjan Balasubramanian, and Aruna Balasubramanian · 2017
Earlier work this paper cites.
Crnn: a joint neural network for redundancy detection
Xinyu Fu, Eugene Ch’ng, Uwe Aickelin, and Simon See · 2017
Earlier work this paper cites.
Deepmon: Mobile gpu-based deep learning framework for continuous vision applications
Loc N Huynh, Youngki Lee, and Rajesh Krishna Balan · 2017
Earlier work this paper cites.
Deepeye: Resource efficient local execution of multiple deep vision models using wearable commodity hardware
Akhil Mathur, Nicholas D Lane, Sourav Bhattacharya, Aidan Boran, Claudio Forlivesi, and Fahim Kawsar · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
On-demand deep model compression for mobile devices: A usage-driven model selection framework
Sicong Liu, Yingyan Lin, Zimu Zhou, Kaiming Nan, Hui Liu, and Junzhao Du · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Deepcache: Principled cache for mobile deep vision
Mengwei Xu, Mengze Zhu, Yunxin Liu, Felix Xiaozhu Lin, and Xuanzhe Liu · 2018
Earlier work this paper cites.
Split-cnn: Splitting window-based operations in convolutional neural networks for memory system optimization
Tian Jin and Seokin Hong · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Deepwear: Adaptive local offloading for on-wearable deep learning
Mengwei Xu, Feng Qian, Mengze Zhu, Feifan Huang, Saumay Pushp, and Xuanzhe Liu · 2019
Earlier work this paper cites.
Compressing neural machine translation models with 4-bit precision
Alham Fikri Aji and Kenneth Heafield · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
Spinn: synergistic progressive inference of neural networks over device and cloud
Stefanos Laskaridis, Stylianos I Venieris, Mario Almeida, Ilias Leontiadis, and Nicholas D Lane · 2020
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2020
Earlier work this paper cites.
Patdnn: Achieving real-time dnn execution on mobile devices with pattern-based weight pruning
Wei Niu, Xiaolong Ma, Sheng Lin, Shihao Wang, Xuehai Qian, Xue Lin, Yanzhi Wang, and Bin Ren · 2020
Earlier work this paper cites.
Gobo: Quantizing attention-based nlp models for low latency and energy efficient inference
Ali Hadi Zadeh, Isak Edo, Omar Mohamed Awad, and Andreas Moshovos · 2020
Earlier work this paper cites.
Base layers: Simplifying training of large, sparse models
Mike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal, and Luke Zettlemoyer · 2021
Earlier work this paper cites.
Enabling large neural networks on tiny microcontrollers with swapping
Hongyu Miao and Felix Xiaozhu Lin · 2021
Earlier work this paper cites.
Hash layers for large sparse models
Stephen Roller, Sainbayar Sukhbaatar, Jason Weston, et al · 2021
Earlier work this paper cites.
https://www.microsoft.com/en-us/research/blog/microsoft-translator-enhanced-with-z-code-mixture-of-experts-models/ , 2022
Microsoft translator enhanced with z-code mixture of experts models · 2022
Earlier work this paper cites.
https://huggingface.co/blog/quanto-introduction , 2022
Quanto: a pytorch quantization backend for optimum · 2022
Earlier work this paper cites.
Ta-moe: Topology-aware large scale mixture-of-expert training
Chang Chen, Min Li, Zhihua Wu, Dianhai Yu, and Chao Yang · 2022
Earlier work this paper cites.
On the representation collapse of sparse mixture of experts
Zewen Chi, Li Dong, Shaohan Huang, Damai Dai, Shuming Ma, Barun Patra, Saksham Singhal, Payal Bajaj, Xia Song, Xian-Ling Mao, et al · 2022
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · 2022
Cited alongside, same era.
Glam: Efficient scaling of language models with mixture-of-experts
Nan Du, Yanping Huang, Andrew M Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, et al · 2022
Cited alongside, same era.
A review of sparse expert models in deep learning
William Fedus, Jeff Dean, and Barret Zoph · 2022
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2022
Cited alongside, same era.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh · 2022
Cited alongside, same era.
Sparsegpt: Massive language models can be accurately pruned in one-shot
Elias Frantar and Dan Alistarh · 2023
Closest in time.
Specializing smaller language models towards multi-step reasoning
Yao Fu, Hao Peng, Litu Ou, Ashish Sabharwal, and Tushar Khot · 2023
Closest in time.
In-context autoencoder for context compression in a large language model
Tao Ge, Jing Hu, Xun Wang, Si-Qing Chen, and Furu Wei · 2023
Closest in time.
Sti: Turbocharge nlp inference at the edge via elastic pipelining
Liwei Guo, Wonkyo Choe, and Felix Xiaozhu Lin · 2023
Closest in time.
Squeezellm: Dense-and-sparse quantization
Sehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong, Xiuyu Li, Sheng Shen, Michael W Mahoney, and Kurt Keutzer · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mixture of quantized experts (moqe): Complementary effect of low-bit quantization and robustness
Young Jin Kim, Raffy Fahim, and Hany Hassan · 2022
Cited alongside, same era.
Romou: Rapidly generate high-performance tensor kernels for mobile gpus
Rendong Liang, Ting Cao, Jicheng Wen, Manni Wang, Yang Wang, Jianhua Zou, and Yunxin Liu · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Melon: Breaking the memory wall for resource-efficient on-device machine learning
Qipeng Wang, Mengwei Xu, Chao Jin, Xinran Dong, Jinliang Yuan, Xin Jin, Gang Huang, Yunxin Liu, and Xuanzhe Liu · 2022
Cited alongside, same era.
Mixture-of-experts with expert choice routing
Yanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du, Yanping Huang, Vincent Zhao, Andrew M Dai, Quoc V Le, James Laudon, et al · 2022
Cited alongside, same era.
https://www.huaweicentral.com/beating-google-and-apple-huawei-brings-large-ai-model-to-mobile-voice-assistant/ , 2023
Beating google and apple, huawei brings large ai model to mobile voice assistant - huawei central · 2023
Cited alongside, same era.
https://huggingface.co/datasets/c4 , 2023
c4 · datasets at hugging face · 2023
Cited alongside, same era.
Serving moe models on resource-constrained edge devices via dynamic expert swapping
Rui Kong, Yuanchun Li, Qingtian Feng, Weijun Wang, Linghe Kong, and Yunxin Liu · 2023
Closest in time.
Symbolic chain-of-thought distillation: Small models can also” think” step-by-step
Liunian Harold Li, Jack Hessel, Youngjae Yu, Xiang Ren, Kai-Wei Chang, and Yejin Choi · 2023
Closest in time.
Losparse: Structured compression of large language models based on low-rank and sparse approximation
Yixiao Li, Yifan Yu, Qingru Zhang, Chen Liang, Pengcheng He, Weizhu Chen, and Tuo Zhao · 2023
Closest in time.
Awq: Activation-aware weight quantization for llm compression and acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han · 2023
Closest in time.
Llm-qat: Data-free quantization aware training for large language models
Zechun Liu, Barlas Oguz, Changsheng Zhao, Ernie Chang, Pierre Stock, Yashar Mehdad, Yangyang Shi, Raghuraman Krishnamoorthi, and Vikas Chandra · 2023
Closest in time.
Deja vu: Contextual sparsity for efficient llms at inference time
Zichang Liu, Jue Wang, Tri Dao, Tianyi Zhou, Binhang Yuan, Zhao Song, Anshumali Shrivastava, Ce Zhang, Yuandong Tian, Christopher Re, et al · 2023
Closest in time.
Llm-pruner: On the structural pruning of large language models
Xinyin Ma, Gongfan Fang, and Xinchao Wang · 2023
Closest in time.
Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Rae Ying Yee Wong, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar, and Zhihao Jia · 2023
Closest in time.
Skeleton-of-thought: Large language models can do parallel decoding
Xuefei Ning, Zinan Lin, Zixuan Zhou, Huazhong Yang, and Yu Wang · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Rishov Sarkar, Hanxue Liang, Zhiwen Fan, Zhangyang Wang, and Cong Hao · 2023
Closest in time.
Accelerating llm inference with staged speculative decoding
Benjamin Spector and Chris Re · 2023
Closest in time.
A simple and effective pruning approach for large language models
Mingjie Sun, Zhuang Liu, Anna Bair, and J Zico Kolter · 2023
Closest in time.
[industry] gkd: A general knowledge distillation framework for large-scale pre-trained language model
Shicheng Tan, Weng Lam Tam, Yuanchun Wang, Wenwen Gong, Shu Zhao, Peng Zhang, and Jie Tang · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
Llmzip: Lossless text compression using large language models
Chandra Shekhara Kaushik Valmeekam, Krishna Narayanan, Dileep Kalathil, Jean-Francois Chamberland, and Srinivas Shakkottai · 2023
Closest in time.
Scott: Self-consistent chain-of-thought distillation
Peifeng Wang, Zhengyang Wang, Zheng Li, Yifan Gao, Bing Yin, and Xiang Ren · 2023
Closest in time.
Lamini-lm: A diverse herd of distilled models from large-scale instructions
Minghao Wu, Abdul Waheed, Chiyu Zhang, Muhammad Abdul-Mageed, and Alham Fikri Aji · 2023
Closest in time.
Smoothquant: Accurate and efficient post-training quantization for large language models
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han · 2023
Closest in time.
Mingxue Xu, Yao Lei Xu, and Danilo P Mandic · 2023
Closest in time.
mllm: fast and lightweight multimodal llm inference engine for mobile and edge devices, 2023
Rongjie Yi, Xiang Li, Zhenyan Lu, Hao Zhang, Daliang Xu, Liming Yang, Weikai Xie, Chenghua Wang, Xuanzhe Liu, and Mengwei Xu · 2023
Closest in time.
Distilling script knowledge from large language models for constrained language planning
Siyu Yuan, Jiangjie Chen, Ziquan Fu, Xuyang Ge, Soham Shah, Charles Robert Jankowski, Deqing Yang, and Yanghua Xiao · 2023
Closest in time.
Rptq: Reorder-based post-training quantization for large language models
Zhihang Yuan, Lin Niu, Jiawei Liu, Wenyu Liu, Xinggang Wang, Yuzhang Shang, Guangyu Sun, Qiang Wu, Jiaxiang Wu, and Bingzhe Wu · 2023
Closest in time.
https://huggingface.co/openbmb/MiniCPM-MoE-8x2B , 2024
Minicpm-moe-8x2b · 2024
Closest in time.
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al · 2024
Closest in time.