Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Optimization of MPI collective communication on BlueGene/L systems. In Proceedings of the 19th annual international conference on Supercomputing
George Almási, Philip Heidelberger, Charles J Archer, Xavier Martorell, C Chris Erway, José E Moreira, Burkhard Steinmacher-Burow, and Yili Zheng. 2005 · 2005
Earlier work this paper cites.
On the feasibility of optical circuit switching for high performance computing systems. In SC’05: Proceedings of the 2005 ACM/IEEE Conference on Supercomputing . 16–16
Kevin J Barker, Alan Benner, Ray Hoare, Adolfy Hoisie, Alex K Jones, Darren K Kerbyson, Dan Li, Rami Melhem, Ramakrishnan Rajamony, Eugen Schenfeld, et al · 2005
Earlier work this paper cites.
Optimization of collective communication operations in MPICH
Rajeev Thakur, Rolf Rabenseifner, and William Gropp. 2005 · 2005
Earlier work this paper cites.
Collective Communication: Theory, Practice, and Experience
Ernie Chan, Marcel Heimlich, Avi Purkayastha, and Robert van de Geijn. 2007 · 2007
Earlier work this paper cites.
Hogwild!: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu. 2011 · 2011
Earlier work this paper cites.
Optical interconnection networks for high-performance computing systems
Aleksandr Biberman and Keren Bergman. 2012 · 2012
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns. In INTERSPEECH . 1058–1062
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu. 2014 · 2014
Earlier work this paper cites.
High throughput data center topology design. In USENIX NSDI . 29–41
Ankit Singla, P Brighten Godfrey, and Alexandra Kolla. 2014 · 2014
Earlier work this paper cites.
Taming the wild: A unified analysis of hogwild-style algorithms
Christopher M De Sa, Ce Zhang, Kunle Olukotun, and Christopher Ré. 2015 · 2015
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Original
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. 2017 · 2017
Earlier work this paper cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Original
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and William J Dally. 2017 · 2017
Earlier work this paper cites.
Device Placement Optimization with Reinforcement Learning. In International Conference on Machine Learning , Vol. 70. 2430–2439
Azalia Mirhoseini, Hieu Pham, Quoc V Le, Benoit Steiner, Rasmus Larsen, Yuefeng Zhou, Naveen Kumar, Mohammad Norouzi, Samy Bengio, and Jeff Dean. 2017 · 2017
Earlier work this paper cites.
Attention Is All You Need
Original
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Wide Residual Networks
Original
Sergey Zagoruyko and Nikos Komodakis. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Original
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Horovod: fast and easy distributed deep learning in TensorFlow
Original
Alexander Sergeev and Mike Del Balso. 2018 · 2018
Earlier work this paper cites.
Mesh-tensorflow: Deep learning for supercomputers
Original
Noam Shazeer, Youlong Cheng, Niki Parmar, Dustin Tran, Ashish Vaswani, Penporn Koanantakool, Peter Hawkins, HyoukJoong Lee, Mingsheng Hong, Cliff Young, et al · 2018
Earlier work this paper cites.
Imagenet training in minutes
Original
Yang You, Zhao Zhang, Cho-Jui Hsieh, James Demmel, and Kurt Keutzer. 2018 · 2018
Earlier work this paper cites.
BlueConnect: Novel hierarchical all-reduce on multi-tired network for deep learning. In Conference on Machine Learning and Systems
Minsik Cho, Ulrich Finkler, and David Kung. 2019 · 2019
Earlier work this paper cites.
Analysis of DAWNBench, a Time-to-Accuracy Machine Learning Performance Benchmark
Cody Coleman, Daniel Kang, Deepak Narayanan, Luigi Nardi, Tian Zhao, Jian Zhang, Peter Bailis, Kunle Olukotun, Chris Ré, and Matei Zaharia. 2019 · 2019
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism. In Advances in Neural Information Processing Systems , Vol. 32
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al · 2019
Earlier work this paper cites.
Priority-based parameter propagation for distributed DNN training
Original
Anand Jayarajan, Jinliang Wei, Garth Gibson, Alexandra Fedorova, and Gennady Pekhimenko. 2019 · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Original
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Earlier work this paper cites.