Fetching the paper…
Reading the bibliography…
Transformer achieves promising results on various tasks.
“Bert: Pre-training of deep bidirectional transformers for language understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2018
Earlier work this paper cites.
“Gpipe: Efficient training of giant neural networks using pipeline parallelism”
Yanping Huang et al · 2018
Earlier work this paper cites.
“High resolution medical image analysis with spatial partitioning”
Le Hou et al · 2019
Earlier work this paper cites.
“BERT with history answer embedding for conversational question answering”
Chen Qu et al · 2019
Earlier work this paper cites.
“Language Models are Unsupervised Multitask Learners”, 2019
Alec Radford et al · 2019
Earlier work this paper cites.
“Megatron-lm: Training multi-billion parameter language models using model parallelism”
Mohammad Shoeybi et al · 2019
Earlier work this paper cites.
“Language models are few-shot learners”
Tom Brown et al · 2020
Earlier work this paper cites.
“An image is worth 16x16 words: Transformers for image recognition at scale”
Alexey Dosovitskiy et al · 2020
Earlier work this paper cites.
Ahmed Elnaggar et al · 2020
Earlier work this paper cites.
“Gshard: Scaling giant models with conditional computation and automatic sharding”
Dmitry Lepikhin et al · 2020
Cited alongside, same era.
“Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer”
Colin Raffel et al · 2020
Cited alongside, same era.
“Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters”
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase and Yuxiong He · 2020
Cited alongside, same era.
“Linformer: Self-attention with linear complexity”
Sinong Wang et al · 2020
Cited alongside, same era.
“An Embarrassingly Simple Model for Dialogue Relation Extraction”
Fuzhao Xue, Aixin Sun, Hao Zhang and Eng Chng · 2020
Cited alongside, same era.
“Document-Level Relation Extraction with Adaptive Thresholding and Localized Context Pooling”
Wenxuan Zhou, Kevin Huang, Tengyu Ma and Jing Huang · 2020
Later among the works it cites.
“Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity”
William Fedus, Barret Zoph and Noam Shazeer · 2021
Closest in time.
“Efficient Large-Scale Language Model Training on GPU Clusters”
Deepak Narayanan et al · 2021
Closest in time.
“Recent advances in deep learning based dialogue systems: A systematic survey”
Jinjie Ni et al · 2021
Closest in time.
“ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“GDPNet: Refining Latent Multi-View Graph for Relation Extraction”
Fuzhao Xue, Aixin Sun, Hao Zhang and Eng Chng · 2020
Cited alongside, same era.
“Bert representations for video question answering”
Zekun Yang et al · 2020
Cited alongside, same era.
“Big bird: Transformers for longer sequences”
Manzil Zaheer et al · 2020
Cited alongside, same era.
“Span-based Localizing Network for Natural Language Video Localization”
Hao Zhang, Aixin Sun, Wei Jing and Joey Zhou · 2020
Cited alongside, same era.
Samyam Rajbhandari et al · 2021
Closest in time.
“ZeRO-Offload: Democratizing Billion-Scale Model Training”, 2021
Jie Ren et al · 2021
Closest in time.
“PSSM-Distil: Protein Secondary Structure Prediction (PSSP) on Low-Quality PSSM by Knowledge Distillation with Contrastive Learning”, 2021
Qin Wang et al · 2021
Closest in time.
“GSPMD: General and Scalable Parallelization for ML Computation Graphs”
Yuanzhong Xu et al · 2021
Closest in time.
“Natural Language Video Localization: A Revisit in Span-based Question Answering Framework”
Hao Zhang et al · 2021
Closest in time.