Fetching the paper…
Reading the bibliography…
By increasing model parameters but activating them sparsely when performing a task, the use of Mixture-of-Experts (MoE) architecture significantly improves the performance of Large Language Models (LLMs) without increasing the inference cost.
Mixtures of expert networks
R Jacobs, MI Jordan, GE Hinton, et al. 1991 · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the em algorithm
Michael I Jordan and Robert A Jacobs. 1994 · 1994
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
On spectral clustering: Analysis and an algorithm
Andrew Ng, Michael Jordan, and Yair Weiss. 2001 · 2001
Earlier work this paper cites.
Empirical studies on the properties of linear regions in deep neural networks
Xiao Zhang and Dongrui Wu. 2020 · 2001
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2020 · 2006
Earlier work this paper cites.
The fifth pascal recognizing textual entailment challenge
Luisa Bentivogli, Peter Clark, Ido Dagan, and Danilo Giampiccolo. 2009 · 2009
Earlier work this paper cites.
Twenty years of mixture of experts
Seniha Esen Yuksel, Joseph N Wilson, and Paul D Gader. 2012 · 2012
Earlier work this paper cites.
Mixture of experts: a literature survey
Saeed Masoudnia and Reza Ebrahimpour. 2014 · 2014
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017 · 2017
Earlier work this paper cites.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Michael Zhu and Suyog Gupta. 2017 · 2017
Earlier work this paper cites.
Rethinking the value of network pruning
Zhuang Liu, Mingjie Sun, Tinghui Zhou, Gao Huang, and Trevor Darrell. 2018 · 2018
Earlier work this paper cites.
Thinet: Pruning cnn filters for a thinner net
Jian-Hao Luo, Hao Zhang, Hong-Yu Zhou, Chen-Wei Xie, Jianxin Wu, and Weiyao Lin. 2018 · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018 · 2018
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Similarity of neural network representations revisited
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton. 2019 · 2019
Cited alongside, same era.
Structured pruning of neural networks with budget-aware regularization
Carl Lemaire, Andrew Achkar, and Pierre-Marc Jodoin. 2019 · 2019
Cited alongside, same era.
Where to prune: Using lstm to guide data-dependent soft pruning
Guiguang Ding, Shuo Zhang, Zizhou Jia, Jing Zhong, and Jungong Han. 2020 · 2020
Cited alongside, same era.
Robust learning with the hilbert-schmidt independence criterion
Daniel Greenfeld and Uri Shalit. 2020 · 2020
Cited alongside, same era.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Mitchell Wortsman, Gabriel Ilharco, Samir Ya Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, et al. 2022 · 2022
Later among the works it cites.
Uni-perceiver-moe: Learning sparse generalist models with conditional moes
Jinguo Zhu, Xizhou Zhu, Wenhai Wang, Xiaohua Wang, Hongsheng Li, Xiaogang Wang, and Jifeng Dai. 2022 · 2022
Later among the works it cites.
St-moe: Designing stable and transferable sparse expert models
Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus. 2022 · 2022
Later among the works it cites.
Depgraph: Towards any structural pruning
Gongfan Fang, Xinyin Ma, Mingli Song, Michael Bi Mi, and Xinchao Wang. 2023 · 2023
Later among the works it cites.
Can unstructured pruning reduce the depth in deep neural networks?
Zhu Liao, Victor Quétu, Van-Tam Nguyen, and Enzo Tartaglione. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Cited alongside, same era.
Pruning from scratch
Yulong Wang, Xiaolu Zhang, Lingxi Xie, Jun Zhou, Hang Su, Bo Zhang, and Xiaolin Hu. 2020 · 2020
Cited alongside, same era.
A gpu architecture aware fine-grain pruning technique for deep neural networks
Kyusik Choi and Hoeseok Yang. 2021 · 2021
Cited alongside, same era.
A framework for few-shot language model evaluation
Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, et al. 2021 · 2021
Cited alongside, same era.
Discrimination-aware network pruning for deep model compression
Jing Liu, Bohan Zhuang, Zhuangwei Zhuang, Yong Guo, Junzhou Huang, Jinhui Zhu, and Mingkui Tan. 2021 · 2021
Cited alongside, same era.
Task-specific expert pruning for sparse mixture-of-experts
Tianyu Chen, Shaohan Huang, Yuan Xie, Binxing Jiao, Daxin Jiang, Haoyi Zhou, Jianxin Li, and Furu Wei. 2022 · 2022
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer. 2022 · 2022
Cited alongside, same era.
Later among the works it cites.
Sheared llama: Accelerating language model pre-training via structured pruning
Mengzhou Xia, Tianyu Gao, Zhiyuan Zeng, and Danqi Chen. 2023 · 2023
Later among the works it cites.
Robust mixture-of-expert training for convolutional neural networks
Yihua Zhang, Ruisi Cai, Tianlong Chen, Guanhua Zhang, Huan Zhang, Pin-Yu Chen, Shiyu Chang, Zhangyang Wang, and Sijia Liu. 2023 · 2023
Later among the works it cites.
Ld-pruner: Efficient pruning of latent diffusion models using task-agnostic insights
Thibault Castells, Hyoung-Kyu Song, Bo-Kyeong Kim, and Shinkook Choi. 2024 · 2024
Closest in time.
A provably effective method for pruning experts in fine-tuned sparse mixture-of-experts
Mohammed Nowaz Rabbani Chowdhury, Meng Wang, Kaoutar El Maghraoui, Naigang Wang, Pin-Yu Chen, and Christopher Carothers. 2024 · 2024
Closest in time.
Demystifying the compression of mixture-of-experts through a unified framework
Shwai He, Daize Dong, Liang Ding, and Ang Li. 2024 · 2024
Closest in time.
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024 · 2024
Closest in time.
Xudong Lu, Qi Liu, Yuhui Xu, Aojun Zhou, Siyuan Huang, Bo Zhang, Junchi Yan, and Hongsheng Li. 2024 · 2024
Closest in time.
What makes a good prune? maximal unstructured pruning for maximal cosine similarity
Gabryel Mason-Williams and Fredrik Dahlqvist. 2024 · 2024
Closest in time.
Towards energy efficient spiking neural networks: An unstructured pruning framework
Xinyu Shi, Jianhao Ding, Zecheng Hao, and Zhaofei Yu. 2024 · 2024
Closest in time.
Model-glue: Democratized llm scaling for a large model zoo in the wild
Xinyu Zhao, Guoheng Sun, Ruisi Cai, Yukun Zhou, Pingzhi Li, Peihao Wang, Bowen Tan, Yexiao He, Li Chen, Yi Liang, Beidi Chen, Binhang Yuan, Hongyi Wang, Ang Li, Zhangyang Wang, and Tianlong Chen. 2024 · 2024
Closest in time.