Fetching the paper…
Reading the bibliography…
Mixture-of-Experts (MoE) models can achieve promising results with outrageous large amount of parameters but constant computation cost, and thus it has become a trend in model scaling.
12-in-1: Multi-task vision and language representation learning
J. Lu, V. Goswami, M. Rohrbach, D. Parikh, and S. Lee · 1912
Earlier work this paper cites.
Oscar: Object-semantics aligned pre-training for vision-language tasks
X. Li, X. Yin, C. Li, P. Zhang, X. Hu, L. Zhang, L. Wang, H. Hu, L. Dong, F. Wei, Y. Choi, and J. Gao · 2004
Earlier work this paper cites.
UNIMO: towards unified-modal understanding and generation via cross-modal contrastive learning
W. Li, C. Gao, G. Niu, X. Xiao, H. Liu, J. Liu, H. Wu, and H. Wang · 2012
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. Kudlur, J. Levenberg, R. Monga, S. Moore, D. G. Murray, B. Steiner, P. A. Tucker, V. Vasudevan, P. Warden, M. Wicke, Y. Yu, and X. Zheng · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Deep contextualized word representations
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
N. Shazeer and M. Stern · 2018
Earlier work this paper cites.
Mesh-tensorflow: Deep learning for supercomputers
N. Shazeer, Y. Cheng, N. Parmar, D. Tran, A. Vaswani, P. Koanantakool, P. Hawkins, H. Lee, M. Hong, C. Young, R. Sepassi, and B. A. Hechtman · 2018
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Unified language model pre-training for natural language understanding and generation
L. Dong, N. Yang, W. Wang, F. Wei, X. Liu, Y. Wang, J. Gao, M. Zhou, and H. Hon · 2019
Earlier work this paper cites.
Roberta: A robustly optimized BERT pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Cited alongside, same era.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
J. Lu, D. Batra, D. Parikh, and S. Lee · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2019
Cited alongside, same era.
Zero: Memory optimization towards training a trillion parameter models
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He · 2019
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Later among the works it cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
D. Lepikhin, H. Lee, Y. Xu, D. Chen, O. Firat, Y. Huang, M. Krikun, N. Shazeer, and Z. Chen · 2020
Later among the works it cites.
Whale: A unified distributed training framework
A. Wang, X. Jia, L. Jiang, J. Zhang, Y. Li, and W. Lin · 2020
Later among the works it cites.
Ernie-vil: Knowledge enhanced vision-language representations through scene graph
F. Yu, J. Tang, W. Yin, Y. Sun, H. Tian, H. Wu, and H. Wang · 2020
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Diversity and depth in per-example routing models
P. Ramachandran and Q. V. Le · 2019
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 2019
Cited alongside, same era.
Vl-bert: Pre-training of generic visual-linguistic representations
W. Su, X. Zhu, Y. Cao, B. Li, L. Lu, F. Wei, and J. Dai · 2019
Cited alongside, same era.
Xlnet: Generalized autoregressive pretraining for language understanding
Z. Yang, Z. Dai, Y. Yang, J. G. Carbonell, R. Salakhutdinov, and Q. V. Le · 2019
Cited alongside, same era.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
UNITER: universal image-text representation learning
Y. Chen, L. Li, L. Yu, A. E. Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu · 2020
Cited alongside, same era.
Large-scale adversarial training for vision-and-language representation learning
Z. Gan, Y.-C. Chen, L. Li, C. Zhu, Y. Cheng, and J. Liu · 2020
Cited alongside, same era.
W. Fedus, B. Zoph, and N. Shazeer · 2021
Closest in time.
Fastmoe: A fast mixture-of-expert training system
J. He, J. Qiu, A. Zeng, Z. Yang, J. Zhai, and J. Tang · 2021
Closest in time.
BASE layers: Simplifying training of large, sparse models
M. Lewis, S. Bhosale, T. Dettmers, N. Goyal, and L. Zettlemoyer · 2021
Closest in time.
M6: A chinese multimodal pretrainer
J. Lin, R. Men, A. Yang, C. Zhou, M. Ding, Y. Zhang, P. Wang, A. Wang, L. Jiang, X. Jia, J. Zhang, J. Zhang, X. Zou, Z. Li, X. Deng, J. Liu, J. Xue, H. Zhou, J. Ma, J. Yu, Y. Li, W. Lin, J. Zhou, J. Tang, and H. Yang · 2021
Closest in time.
Zero-infinity: Breaking the GPU memory wall for extreme scale deep learning
S. Rajbhandari, O. Ruwase, J. Rasley, S. Smith, and Y. He · 2021
Closest in time.
Zero-shot text-to-image generation, 2021
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Closest in time.
Zero-offload: Democratizing billion-scale model training
J. Ren, S. Rajbhandari, R. Y. Aminabadi, O. Ruwase, S. Yang, M. Zhang, D. Li, and Y. He · 2021
Closest in time.
Vinvl: Making visual representations matter in vision-language models
P. Zhang, X. Li, X. Hu, J. Yang, L. Zhang, L. Wang, Y. Choi, and J. Gao · 2021
Closest in time.