Fetching the paper…
Reading the bibliography…
More transformer blocks with residual connections have recently achieved impressive results on various tasks.
Reducing BERT pre-training time from 3 days to 76 minutes
You, Y.; Li, J.; Hseu, J.; Song, X.; Demmel, J.; and Hsieh, C.-J. 2019a · 1904
Earlier work this paper cites.
Large batch optimization for deep learning: Training bert in 76 minutes
You, Y.; Li, J.; Reddi, S.; Hseu, J.; Kumar, S.; Bhojanapalli, S.; Song, X.; Demmel, J.; Keutzer, K.; and Hsieh, C.-J. 2019b · 1904
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Lan, Z.; Chen, M.; Goodman, S.; Gimpel, K.; Sharma, P.; and Soricut, R. 2019 · 1909
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M.; Patwary, M.; Puri, R.; LeGresley, P.; Casper, J.; and Catanzaro, B. 2019 · 1909
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2019 · 1910
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D.; Lee, H.; Xu, Y.; Chen, D.; Firat, O.; Huang, Y.; Krikun, M.; Shazeer, N.; and Chen, Z. 2020 · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A.; Hinton, G.; et al. 2009 · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
Document-Level Relation Extraction with Adaptive Thresholding and Localized Context Pooling
Zhou, W.; Huang, K.; Ma, T.; and Huang, J. 2020 · 2010
Earlier work this paper cites.
An Embarrassingly Simple Model for Dialogue Relation Extraction
Xue, F.; Sun, A.; Zhang, H.; and Chng, E. S. 2020a · 2012
Earlier work this paper cites.
GDPNet: Refining Latent Multi-View Graph for Relation Extraction
Xue, F.; Sun, A.; Zhang, H.; and Chng, E. S. 2020b · 2012
Earlier work this paper cites.
Deep learning of representations: Looking forward
Bengio, Y. 2013 · 2013
Cited alongside, same era.
Zhu, Y.; Kiros, R.; Zemel, R.; Salakhutdinov, R.; Urtasun, R.; Torralba, A.; and Fidler, S. 2015 · 2015
Cited alongside, same era.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Rajpurkar, P.; Zhang, J.; Lopyrev, K.; and Liang, P. 2016 · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Szegedy, C.; Vanhoucke, V.; Ioffe, S.; Shlens, J.; and Wojna, Z. 2016 · 2016
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Later among the works it cites.
BERT with history answer embedding for conversational question answering
Qu, C.; Yang, L.; Qiu, M.; Croft, W. B.; Zhang, Y.; and Iyyer, M. 2019 · 2019
Later among the works it cites.
Patient Knowledge Distillation for BERT Model Compression
Sun, S.; Cheng, Y.; Gan, Z.; and Liu, J. 2019 · 2019
Later among the works it cites.
Randaugment: Practical automated data augmentation with a reduced search space
Cubuk, E. D.; Zoph, B.; Shlens, J.; and Le, Q. V. 2020 · 2020
Later among the works it cites.
BERT-of-Theseus: Compressing BERT by Progressive Module Replacing
Xu, C.; Zhou, W.; Ge, T.; Wei, F.; and Zhou, M. 2020 · 2020
Later among the works it cites.
Bert representations for video question answering
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shazeer, N.; Mirhoseini, A.; Maziarz, K.; Davis, A.; Le, Q.; Hinton, G.; and Dean, J. 2017 · 2017
Cited alongside, same era.
Revisiting unreasonable effectiveness of data in deep learning era
Sun, C.; Shrivastava, A.; Singh, S.; and Gupta, A. 2017 · 2017
Cited alongside, same era.
mixup: Beyond empirical risk minimization
Zhang, H.; Cisse, M.; Dauphin, Y. N.; and Lopez-Paz, D. 2017 · 2017
Cited alongside, same era.
Dehghani, M.; Gouws, S.; Vinyals, O.; Uszkoreit, J.; and Kaiser, Ł. 2018 · 2018
Cited alongside, same era.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y.; Cheng, Y.; Bapna, A.; Firat, O.; Chen, M. X.; Chen, D.; Lee, H.; Ngiam, J.; Le, Q. V.; Wu, Y.; et al. 2018 · 2018
Cited alongside, same era.
Know What You Don’t Know: Unanswerable Questions for SQuAD
Rajpurkar, P.; Jia, R.; and Liang, P. 2018 · 2018
Cited alongside, same era.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. 2018 · 2018
Cited alongside, same era.
Yang, Z.; Garcia, N.; Chu, C.; Otani, M.; Nakashima, Y.; and Takemura, H. 2020 · 2020
Later among the works it cites.
Revisiting knowledge distillation via label smoothing regularization
Yuan, L.; Tay, F. E.; Li, G.; Wang, T.; and Feng, J. 2020 · 2020
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W.; Zoph, B.; and Shazeer, N. 2021 · 2021
Closest in time.
Sequence Parallelism: Making 4D Parallelism Possible
Li, S.; Xue, F.; Li, Y.; and You, Y. 2021 · 2021
Closest in time.
Scaling Vision with Sparse Mixture of Experts
Riquelme, C.; Puigcerver, J.; Mustafa, B.; Neumann, M.; Jenatton, R.; Pinto, A. S.; Keysers, D.; and Houlsby, N. 2021 · 2021
Closest in time.
Exploring Sparse Expert Models and Beyond
Yang, A.; Lin, J.; Men, R.; Zhou, C.; Jiang, L.; Jia, X.; Wang, A.; Zhang, J.; Wang, J.; Li, Y.; et al. 2021 · 2021
Closest in time.
Zhai, X.; Kolesnikov, A.; Houlsby, N.; and Beyer, L. 2021 · 2021
Closest in time.