Fetching the paper…
Reading the bibliography…
Larger transformer models always perform better on various tasks but require more costs to scale up the model size.
Zipf’s law everywhere
Wentian Li · 2002
Earlier work this paper cites.
On reinforcement learning for deep neural architectures: conditional computation with stochastic computation policies
Emmanuel Bengio · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Routing networks and the challenges of modular and compositional computation
Clemens Rosenbaum, Ignacio Cases, Matthew Riemer, and Tim Klinger · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Don’t waste your bits! squeeze activations and gradients for deep neural networks via tinyscript
Fangcheng Fu, Yuzheng Hu, Yihan He, Jiawei Jiang, Yingxia Shao, Ce Zhang, and Bin Cui · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Reformer: The efficient transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya · 2020
Earlier work this paper cites.
Efficient large scale language modeling with mixtures of experts
Mikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov, Myle Ott, Sam Shleifer, Xi Victoria Lin, Jingfei Du, Srinivasan Iyer, Ramakanth Pasunuru, et al · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2021
Earlier work this paper cites.
M6: A chinese multimodal pretrainer
Junyang Lin, Rui Men, An Yang, Chang Zhou, Ming Ding, Yichang Zhang, Peng Wang, Ang Wang, Le Jiang, Xianyan Jia, et al · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Cited alongside, same era.
Evomoe: An evolutional mixture-of-experts training framework via dense-to-sparse gate
Xiaonan Nie, Xupeng Miao, Shijie Cao, Lingxiao Ma, Qibin Liu, Jilong Xue, Youshan Miao, Yi Liu, Zhi Yang, and Bin Cui · 2021
Cited alongside, same era.
Scaling vision with sparse mixture of experts
Carlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann, Rodolphe Jenatton, André Susano Pinto, Daniel Keysers, and Neil Houlsby · 2021
Cited alongside, same era.
Hash layers for large sparse models
Stephen Roller, Sainbayar Sukhbaatar, Jason Weston, et al · 2021
Cited alongside, same era.
Accelerate distributed deep learning with cluster-aware sketch quantization
Keshi Ge, Yiming Zhang, Yongquan Fu, Zhiquan Lai, Xiaoge Deng, and Dongsheng Li · 2023
Later among the works it cites.
Tutel: Adaptive mixture-of-experts at scale
Changho Hwang, Wei Cui, Yifan Xiong, Ziyue Yang, Ze Liu, Han Hu, Zilong Wang, Rafael Salas, Jithin Jose, Prabhat Ram, et al · 2023
Later among the works it cites.
Pre-gated moe: An algorithm-system co-design for fast and scalable mixture-of-expert inference
Ranggi Hwang, Jianyu Wei, Shijie Cao, Changho Hwang, Xiaohu Tang, Ting Cao, Mao Yang, and Minsoo Rhu · 2023
Later among the works it cites.
Hetu: A highly efficient automatic parallel distributed deep learning system
Xupeng Miao, Xiaonan Nie, Hailin Zhang, Tong Zhao, and Bin Cui · 2023
Later among the works it cites.
Angel-ptm: A scalable and economical large-scale pre-training system in tencent
Xiaonan Nie, Yi Liu, Fangcheng Fu, Jinbao Xue, Dian Jiao, Xupeng Miao, Yangyu Tao, and Bin Cui · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M6-t: Exploring sparse expert models and beyond
An Yang, Junyang Lin, Rui Men, Chang Zhou, Le Jiang, Xianyan Jia, Ang Wang, Jie Zhang, Jiamang Wang, Yong Li, et al · 2021
Cited alongside, same era.
Ta-moe: Topology-aware large scale mixture-of-expert training
Chang Chen, Min Li, Zhihua Wu, Dianhai Yu, and Chao Yang · 2022
Cited alongside, same era.
Unified scaling laws for routed language models
Aidan Clark, Diego De Las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake Hechtman, Trevor Cai, Sebastian Borgeaud, et al · 2022
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2022
Cited alongside, same era.
Multimodal contrastive learning with limoe: the language-image mixture of experts
Basil Mustafa, Carlos Riquelme, Joan Puigcerver, Rodolphe Jenatton, and Neil Houlsby · 2022
Cited alongside, same era.
Hetumoe: An efficient trillion-scale mixture-of-expert distributed training system
Xiaonan Nie, Pinxue Zhao, Xupeng Miao, Tong Zhao, and Bin Cui · 2022
Cited alongside, same era.
DeepSpeed-MoE: Advancing mixture-of-experts inference and training to power next-generation AI scale
Samyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang, Reza Yazdani Aminabadi, Ammar Ahmad Awan, Jeff Rasley, and Yuxiong He · 2022
Cited alongside, same era.
Later among the works it cites.
Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement
Xiaonan Nie, Xupeng Miao, Zilong Wang, Zichao Yang, Jilong Xue, Lingxiao Ma, Gang Cao, and Bin Cui · 2023
Later among the works it cites.
A communication efficient admm-based distributed algorithm using two-dimensional torus grouping allreduce
Guozheng Wang, Yongmei Lei, Zeyu Zhang, and Cunlu Peng · 2023
Later among the works it cites.
xccl: A survey of industry-led collective communication libraries for deep learning
Adam Weingram, Yuke Li, Hao Qi, Darren Ng, Liuyao Dai, and Xiaoyi Lu · 2023
Later among the works it cites.
Combining graph contrastive embedding and multi-head cross-attention transfer for cross-domain recommendation
Shuo Xiao, Dongqing Zhu, Chaogang Tang, and Zhenzhen Huang · 2023
Later among the works it cites.
Improving BERT fine-tuning via self-ensemble and self-distillation
Yige Xu, Xipeng Qiu, Ligao Zhou, and Xuanjing Huang · 2023
Later among the works it cites.
Scomoe: Efficient mixtures of experts with structured communication
Zhiyuan Zeng and Deyi Xiong · 2023
Later among the works it cites.
Mixtral of experts, 2024
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2024
Closest in time.
How to estimate the number of parameters in transformer models
Dmytro Nikolaiev · 2024
Closest in time.