Fetching the paper…
Reading the bibliography…
Codistillation has been proposed as a mechanism to share knowledge among concurrently trained models by encouraging them to represent the same function through an auxiliary loss.
Learning multiple layers of features from tiny images
A. Krizhevsky, G. Hinton, et al · 2009
Earlier work this paper cites.
International workshop on spoken language translation
M. de la Chimie · 2010
Earlier work this paper cites.
Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning
Z. Allen-Zhu and Y. Li · 2012
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Earlier work this paper cites.
Deep learning , volume 1
I. Goodfellow, Y. Bengio, A. Courville, and Y. Bengio · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Earlier work this paper cites.
Automatic differentiation in pytorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Large scale distributed neural network training through online distillation
R. Anil, G. Pereyra, A. Passos, R. Ormandi, G. E. Dahl, and G. E. Hinton · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Scaling neural machine translation
M. Ott, S. Edunov, D. Grangier, and M. Auli · 2018
Cited alongside, same era.
Bag of tricks for image classification with convolutional neural networks
T. He, Z. Zhang, H. Zhang, Z. Zhang, J. Xie, and M. Li · 2019
Later among the works it cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Y. Huang, Y. Cheng, A. Bapna, O. Firat, D. Chen, M. Chen, H. Lee, J. Ngiam, Q. V. Le, Y. Wu, et al · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli · 2019
Later among the works it cites.
Megatron-LM: Training multi-billion parameter language models using GPU model parallelism
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 2019
Later among the works it cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. U. Stich · 2018
Cited alongside, same era.
Gradient diversity: A key ingredient for scalable distributed learning
D. Yin, A. Pananjady, M. Lam, D. Papailiopoulos, K. Ramchandran, and P. Bartlett · 2018
Cited alongside, same era.
Deep mutual learning
Y. Zhang, T. Xiang, T. M. Hospedales, and H. Lu · 2018
Cited alongside, same era.
Stochastic gradient push for distributed deep learning
M. Assran, N. Loizou, N. Ballas, and M. Rabbat · 2019
Cited alongside, same era.
Closest in time.
Towards the systematic reporting of the energy and carbon footprints of machine learning
P. Henderson, J. Hu, J. Romoff, E. Brunskill, D. Jurafsky, and J. Pineau · 2020
Closest in time.
AdaScale SGD: A user-friendly algorithm for distributed training
T. B. Johnson, P. Agrawal, H. Gu, and C. Guestrin · 2020
Closest in time.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Closest in time.
Gshard: Scaling giant models with conditional computation and automatic sharding
D. Lepikhin, H. Lee, Y. Xu, D. Chen, O. Firat, Y. Huang, M. Krikun, N. Shazeer, and Z. Chen · 2020
Closest in time.