Fetching the paper…
Reading the bibliography…
The ever-growing model size and scale of compute have attracted increasing interests in training deep learning models over multiple nodes.
Parallel data, tools and interfaces in opus
Tiedemann, J · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P., J. Ba · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., G. Hinton, A. Krizhevsky, et al · 2014
Earlier work this paper cites.
Managed communication and consistency for fast data-parallel iterative analytics
Wei, J., W. Dai, A. Qiao, et al · 2015
Earlier work this paper cites.
8-bit approximations for parallelism in deep learning
Dettmers, T · 2015
Earlier work this paper cites.
Geeps: Scalable deep learning on distributed gpus with a gpu-specialized parameter server
Cui, H., H. Zhang, G. R. Ganger, et al · 2016
Earlier work this paper cites.
The united nations parallel corpus v1. 0
Ziemski, M., M. Junczys-Dowmunt, B. Pouliquen · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks, D., K. Gimpel · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C., V. Vanhoucke, S. Ioffe, et al · 2016
Earlier work this paper cites.
Poseidon: An efficient communication architecture for distributed deep learning on { \{ GPU } \} clusters
Zhang, H., Z. Zheng, S. Xu, et al · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., N. Shazeer, N. Parmar, et al · 2017
Earlier work this paper cites.
Sparse communication for distributed gradient descent
Aji, A. F., K. Heafield · 2017
Earlier work this paper cites.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Alistarh, D., D. Grubic, J. Li, et al · 2017
Earlier work this paper cites.
Literature survey on low rank approximation of matrices
Kishore Kumar, N., J. Schneider · 2017
Earlier work this paper cites.
The iit bombay english-hindi parallel corpus
Kunchukuttan, A., P. Mehta, P. Bhattacharyya · 2017
Earlier work this paper cites.
Mesh-tensorflow: Deep learning for supercomputers
Shazeer, N., Y. Cheng, N. Parmar, et al · 2018
Earlier work this paper cites.
Pipedream: Fast and efficient pipeline parallel dnn training
Harlap, A., D. Narayanan, A. Phanishayee, et al · 2018
Earlier work this paper cites.
Atomo: Communication-efficient learning via atomic sparsification
Wang, H., S. Sievert, S. Liu, et al · 2018
Earlier work this paper cites.
Protection against reconstruction and its applications in private federated learning
Bhowmick, A., J. Duchi, J. Freudiger, et al · 2018
Earlier work this paper cites.
Privacy-preserving deep learning via additively homomorphic encryption
Phong, L. T., Y. Aono, T. Hayashi, et al · 2018
Earlier work this paper cites.
Xnli: Evaluating cross-lingual sentence representations
Conneau, A., G. Lample, R. Rinott, et al · 2018
Cited alongside, same era.
Toward understanding the impact of staleness in distributed machine learning
Dai, W., Y. Zhou, N. Dong, et al · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., M.-W. Chang, K. Lee, et al · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., J. Wu, R. Child, et al · 2019
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M., M. Patwary, R. Puri, et al · 2019
Cited alongside, same era.
K-adapter: Infusing knowledge into pre-trained models with adapters
Wang, R., D. Tang, N. Duan, et al · 2020
Later among the works it cites.
Infoxlm: An information-theoretic framework for cross-lingual language model pre-training
Chi, Z., L. Dong, F. Wei, et al · 2020
Later among the works it cites.
Veco: Variable and flexible cross-lingual pre-training for language understanding and generation
Luo, F., W. Wang, J. Liu, et al · 2020
Later among the works it cites.
Cpm-2: Large-scale cost-effective pre-trained language models
Zhang, Z., Y. Gu, X. Han, et al · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Powersgd: Practical low-rank gradient compression for distributed optimization
Vogels, T., S. P. Karimireddy, M. Jaggi · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for nlp
Houlsby, N., A. Giurgiu, S. Jastrzebski, et al · 2019
Cited alongside, same era.
Ccnet: Extracting high quality monolingual datasets from web crawl data
Wenzek, G., M.-A. Lachaux, A. Conneau, et al · 2019
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Conneau, A., K. Khandelwal, N. Goyal, et al · 2019
Cited alongside, same era.
Wikimatrix: Mining 135m parallel sentences in 1620 language pairs from wikipedia
Schwenk, H., V. Chaudhary, S. Sun, et al · 2019
Cited alongside, same era.
Cross-lingual language model pretraining
Lample, G., A. Conneau · 2019
Cited alongside, same era.
Unicoder: A universal language encoder by pre-training with multiple cross-lingual tasks
Huang, H., Y. Liang, N. Duan, et al · 2019
Cited alongside, same era.
Zeng, W., X. Ren, T. Su, et al · 2021
Later among the works it cites.
Wang, S., Y. Sun, Y. Xiang, et al · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
Rae, J. W., S. Borgeaud, T. Cai, et al · 2021
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., B. Zoph, N. Shazeer · 2021
Later among the works it cites.
M6: A chinese multimodal pretrainer
Lin, J., R. Men, A. Yang, et al · 2021
Later among the works it cites.
Memory-efficient pipeline-parallel dnn training
Narayanan, D., A. Phanishayee, K. Shi, et al · 2021
Later among the works it cites.
Towards scalable distributed training of deep learning on public cloud clusters
Shi, S., X. Zhou, S. Song, et al · 2021
Later among the works it cites.
Msp: Multi-stage prompting for making pre-trained language models better translators
Tan, Z., X. Zhang, S. Wang, et al · 2021
Later among the works it cites.
Xlm-e: cross-lingual language model pre-training via electra
Chi, Z., S. Huang, L. Dong, et al · 2021
Later among the works it cites.
Beyond english-centric multilingual machine translation
Fan, A., S. Bhosale, H. Schwenk, et al · 2021
Later among the works it cites.
Coco-lm: Correcting and contrasting text sequences for language model pretraining
Meng, Y., C. Xiong, P. Bajaj, et al · 2021
Later among the works it cites.
End-to-end adaptive distributed training on paddlepaddle
Ao, Y., Z. Wu, D. Yu, et al · 2021
Later among the works it cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., S. Narang, J. Devlin, et al · 2022
Closest in time.
Federated learning challenges and opportunities: An outlook
Ding, J., E. Tramel, A. K. Sahu, et al · 2022
Closest in time.
Pathways: Asynchronous distributed dataflow for ml
Barham, P., A. Chowdhery, J. Dean, et al · 2022
Closest in time.