Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) typically generate outputs token by token using a fixed compute budget, leading to inefficient resource utilization.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Kelsey Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Once-for-all: Train one network and specialize it for efficient deployment, 2019
Han Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang, and Song Han · 2019
Earlier work this paper cites.
What does bert look at? an analysis of bert’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Conditional computation for continual learning, 2019
Min Lin, Jie Fu, and Yoshua Bengio · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2019
Earlier work this paper cites.
Analyzing the structure of attention in a transformer language model
Jesse Vig and Yonatan Belinkov · 2019
Earlier work this paper cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Miles Subbiah, Jared Kaplan, Prafulla Dhariwal, et al · 2020
Earlier work this paper cites.
Analyzing individual neurons in pre-trained language models, 2020
Nadir Durrani, Hassan Sajjad, Fahim Dalvi, and Yonatan Belinkov · 2020
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories, 2020
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy · 2020
Earlier work this paper cites.
Scalable transfer learning with expert models, 2020
Joan Puigcerver, Carlos Riquelme, Basil Mustafa, Cedric Renggli, André Susano Pinto, Sylvain Gelly, Daniel Keysers, and Neil Houlsby · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Chris Shinn, Adam Roberts, Kenton Lee, Shinn Narang, Michael Matena, et al · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber · 2020
Earlier work this paper cites.
Understanding dataset difficulty with 𝒱 \mathcal{V} -usable information, 2021
Kawin Ethayarajh, Yejin Choi, and Swabha Swayamdipta · 2021
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2021
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Cited alongside, same era.
Megablocks: Efficient sparse training with mixture-of-experts, 2022
Trevor Gale, Deepak Narayanan, Cliff Young, and Matei Zaharia · 2022
Cited alongside, same era.
Fast inference from transformers via speculative decoding, 2022
Yaniv Leviathan, Matan Kalman, and Yossi Matias · 2022
Cited alongside, same era.
The lazy neuron phenomenon: On emergence of activation sparsity in transformers, 2022
Zonglin Li, Chong You, Srinadh Bhojanapalli, Daliang Li, Ankit Singh Rawat, Sashank J. Reddi, Ke Ye, Felix Chern, Felix Yu, Ruiqi Guo, and Sanjiv Kumar · 2022
Cited alongside, same era.
Layerskip: Enabling early exit inference and self-speculative decoding, 2024
Mostafa Elhoushi, Akshat Shrivastava, Diana Liskovich, Basil Hosmer, Bram Wasti, Liangzhen Lai, Anas Mahmoud, Bilge Acun, Saurabh Agarwal, Ahmed Roman, Ahmed A Aly, Beidi Chen, and Carole-Jean Wu · 2024
Closest in time.
Mixture-of-modules: Reinventing transformers as dynamic assemblies of modules
Zhuocheng Gong, Ang Lv, Jian Guan, Junxi Yan, Wei Wu, Huishuai Zhang, Minlie Huang, Dongyan Zhao, and Rui Yan · 2024
Closest in time.
Codeparrot: Github code dataset
Hugging Face · 2024
Closest in time.
Moma: Efficient early-fusion pre-training with mixture of modality-aware experts
Xi Victoria Lin, Akshat Shrivastava, Liang Luo, Srinivasan Iyer, Mike Lewis, Gargi Gosh, Luke S. Zettlemoyer, and Armen Aghajanyan · 2024
Closest in time.
Starcoder 2 and the stack v2: The next generation, 2024
Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, Tianyang Liu, Max Tian, Denis Kocetkov, Arthur Zucker, Younes Belkada, Zijian Wang, Qian Liu, Dmitry Abulkhanov, Indraneil Paul, Zhuang Li, Wen-Ding Li, Megan Risdal, Jia Li, Jian Zhu, Terry Yue Zhuo, Evgenii Zheltonozhskii, Nii Osae Osae Dade, Wenhao Yu, Lucas Krauß, Naman Jain, Yixuan Su, Xuanli He, Manan Dey, Edoardo Abati, Yekun Chai, Niklas Muennighoff, Xiangru Tang, Muhtasham Oblokulov, Christopher Akiki, Marc Marone, Chenghao Mou, Mayank Mishra, Alex Gu, Binyuan Hui, Tri Dao, Armel Zebaze, Olivier Dehaene, Nicolas Patry, Canwen Xu, Julian McAuley, Han Hu, Torsten Scholak, Sebastien Paquet, Jennifer Robinson, Carolyn Jane Anderson, Nicolas Chapados, Mostofa Patwary, Nima Tajbakhsh, Yacine Jernite, Carlos Muñoz Ferrandis, Lingming Zhang, Sean Hughes, Thomas Wolf, Arjun Guha, Leandro von Werra, and Harm de Vries · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mixture-of-experts with expert choice routing, 2022
Yanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du, Yanping Huang, Vincent Zhao, Andrew Dai, Zhifeng Chen, Quoc Le, and James Laudon · 2022
Cited alongside, same era.
St-moe: Designing stable and transferable sparse expert models, 2022
Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus · 2022
Cited alongside, same era.
Skipdecode: Autoregressive skip decoding with batching and caching for efficient llm inference, 2023
Luciano Del Corro, Allie Del Giorno, Sahaj Agarwal, Bin Yu, Ahmed Awadallah, and Subhabrata Mukherjee · 2023
Cited alongside, same era.
Matformer: Nested transformer for elastic inference, 2023
Devvrit, Sneha Kudugunta, Aditya Kusupati, Tim Dettmers, Kaifeng Chen, Inderjit Dhillon, Yulia Tsvetkov, Hannaneh Hajishirzi, Sham Kakade, Ali Farhadi, and Prateek Jain · 2023
Cited alongside, same era.
Orchestrallm: Efficient orchestration of language models for dialogue state tracking
Chia-Hsuan Lee, Hao Cheng, and Mari Ostendorf · 2023
Cited alongside, same era.
Sharcs: Efficient transformers through routing with dynamic width sub-networks, 2023
Mohammadreza Salehi, Sachin Mehta, Aditya Kusupati, Ali Farhadi, and Hannaneh Hajishirzi · 2023
Cited alongside, same era.
Joma: Demystifying multilayer transformers via joint dynamics of mlp and attention, 2023
Yuandong Tian, Yiping Wang, Zhenyu Zhang, Beidi Chen, and Simon Du · 2023
Cited alongside, same era.
Closest in time.
Not all experts are equal: Efficient expert pruning and skipping for mixture-of-experts large language models, 2024
Xudong Lu, Qi Liu, Yuhui Xu, Aojun Zhou, Siyuan Huang, Bo Zhang, Junchi Yan, and Hongsheng Li · 2024
Closest in time.
Ee-tuning: An economical yet scalable solution for tuning early-exit large language models, 2024
Xuchen Pan, Yanxi Chen, Yaliang Li, Bolin Ding, and Jingren Zhou · 2024
Closest in time.
The fineweb datasets: Decanting the web for the finest text data at scale, 2024
Guilherme Penedo, Hynek Kydlíček, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf · 2024
Closest in time.
Mixture-of-depths: Dynamically allocating compute in transformer-based language models, 2024
David Raposo, Sam Ritter, Blake Richards, Timothy Lillicrap, Peter Conway Humphreys, and Adam Santoro · 2024
Closest in time.
Dolma: an open corpus of three trillion tokens for language model pretraining research, 2024
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh, Luke Zettlemoyer, Noah A. Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, and Kyle Lo · 2024
Closest in time.
Hmoe: Heterogeneous mixture of experts for language modeling, 2024
An Wang, Xingwu Sun, Ruobing Xie, Shuaipeng Li, Jiaqi Zhu, Zhen Yang, Pinxue Zhao, J. N. Han, Zhanhui Kang, Di Wang, Naoaki Okazaki, and Cheng zhong Xu · 2024
Closest in time.
Do llamas work in english? on the latent language of multilingual transformers, 2024
Chris Wendler, Veniamin Veselovsky, Giovanni Monea, and Robert West · 2024
Closest in time.
Routing experts: Learning to route dynamic experts in multi-modal large language models
Qiong Wu, Zhaoxi Ke, Yiyi Zhou, Gen Luo, Xiaoshuai Sun, and Rongrong Ji · 2024
Closest in time.
Toward inference-optimal mixture-of-expert large language models, 2024
Longfei Yun, Yonghao Zhuang, Yao Fu, Eric P Xing, and Hao Zhang · 2024
Closest in time.