Fetching the paper…
Reading the bibliography…
Recent results in language understanding using neural networks have required training hardware of unprecedentedscale, with thousands of chips cooperating on a single training run.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Large scale distributed deep networks
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., Ranzato, M., Senior, A., Tucker, P., Yang, K., et al · 2012
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Addressing the rare word problem in neural machine translation
Luong, M.-T., Sutskever, I., Le, Q. V., Vinyals, O., and Zaremba, W · 2014
Earlier work this paper cites.
Criteo releases industry’s largest-ever dataset for machine learning to academic community
Ferns, E · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al · 2015
Earlier work this paper cites.
TensorFlow: A system for large-scale machine learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., et al · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., et al · 2016
Earlier work this paper cites.
Google Vizier: A service for black-box optimization
Golovin, D., Solnik, B., Moitra, S., Kochanski, G., Karro, J., and Sculley, D · 2017
Cited alongside, same era.
Accurate, large minibatch SGD: Training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Cited alongside, same era.
In-datacenter performance analysis of a tensor processing unit
Jouppi, N. P., Young, C., Patil, N., Patterson, D., Agrawal, G., Bajwa, R., Bates, S., Bhatia, S., Boden, N., Borchers, A., et al · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Large batch training of convolutional networks
You, Y., Gitman, I., and Ginsburg, B · 2017
Cited alongside, same era.
Mattson, P., Cheng, C., Coleman, C., Diamos, G., Micikevicius, P., Patterson, D., Tang, H., Wei, G.-Y., Bailis, P., Bittorf, V., et al · 2019
Later among the works it cites.
Deep learning recommendation model for personalization and recommendation systems
Naumov, M., Mudigere, D., Shi, H.-J. M., Huang, J., Sundaraman, N., Park, J., Wang, X., Gupta, U., Wu, C.-J., Azzolini, A. G., et al · 2019
Later among the works it cites.
BFloat16: The secret to high performance on Cloud TPUs
Wang, S. and Kanwar, P · 2019
Later among the works it cites.
Large batch optimization for deep learning: Training BERT in 76 minutes
You, Y., Li, J., Reddi, S., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X. d., Demmel, J., Keutzer, K., and Hsieh, C.-J · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
AI and compute
Amodei, D., Hernandez, D., SastryJack, G., Brockman, C., and Sutskever, I · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Compiling machine learning programs via high-level tracing
Frostig, R., Johnson, M. J., and Leary, C · 2018
Cited alongside, same era.
Measuring the effects of data parallelism on neural network training
Shallue, C. J., Lee, J., Antognini, J., Sohl-Dickstein, J., Frostig, R., and Dahl, G. E · 2018
Cited alongside, same era.
Mesh-TensorFlow: Deep learning for supercomputers
Shazeer, N., Cheng, Y., Parmar, N., Tran, D., Vaswani, A., Koanantakool, P., Hawkins, P., Lee, H., Hong, M., Young, C., et al · 2018
Cited alongside, same era.
Scale MLPerf-0.6 models on Google TPU-v3 pods
Kumar, S., Bitorff, V., Chen, D., Chou, C., Hechtman, B., Lee, H., Kumar, N., Mattson, P., Wang, S., Wang, T., et al · 2019
Cited alongside, same era.
https://mlcommons.org/
MLPerf results used: 0.7-1,17-56,64-70. MLPerf is a trademark of mlcommons.org
Cited in the paper.
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Closest in time.
A domain-specific supercomputer for training deep neural networks
Jouppi, N. P., Yoon, D. H., Kurian, G., Li, S., Patil, N., Laudon, J., Young, C., and Patterson, D · 2020
Closest in time.
Microsoft announces new supercomputer, lays out vision for future AI work
Langston, J · 2020
Closest in time.
Gshard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2020
Closest in time.
NVIDIA Revenue 2006-2020
MacroTrends.net · 2020
Closest in time.
XLA: Optimizing compiler for machine learning
TensorFlow.org · 2020
Closest in time.
Automatic cross-replica sharding of weight update in data-parallel training
Xu, Y., Lee, H., Chen, D., Choi, H., Hechtman, B., and Wang, S · 2020
Closest in time.