Fetching the paper…
Reading the bibliography…
We present the design of a new large scale orchestration layer for accelerators.
Gang scheduling performance benefits for fine-grain synchronization
Dror G. Feitelson and Larry Rudolph · 1992
Earlier work this paper cites.
The MPI message passing interface standard
Lyndon Clarke, Ian Glendinning, and Rolf Hempel · 1994
Earlier work this paper cites.
NVIDIA CUDA software and GPU parallel computing architecture
David Kirk · 2007
Earlier work this paper cites.
The multikernel: A new OS architecture for scalable multicore systems
Andrew Baumann, Paul Barham, Pierre-Evariste Dagand, Tim Harris, Rebecca Isaacs, Simon Peter, Timothy Roscoe, Adrian Schüpbach, and Akhilesh Singhania · 2009
Earlier work this paper cites.
An operating system for multicore and clouds: Mechanisms and implementation
David Wentzlaff, Charles Gruenwald III, Nathan Beckmann, Kevin Modzelewski, Adam Belay, Lamia Youseff, Jason Miller, and Anant Agarwal · 2010
Earlier work this paper cites.
Pegasus: Coordinated scheduling for virtualized accelerator-based systems
Vishakha Gupta, Karsten Schwan, Niraj Tolia, Vanish Talwar, and Parthasarathy Ranganathan · 2011
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
MillWheel: Fault-tolerant stream processing at internet scale
Tyler Akidau, Alex Balikov, Kaya Bekiroğlu, Slava Chernyak, Josh Haberman, Reuven Lax, Sam McVeety, Daniel Mills, Paul Nordstrom, and Sam Whittle · 2013
Earlier work this paper cites.
Naiad: A timely dataflow system
Derek Murray, Frank McSherry, Rebecca Isaacs, Michael Isard, Paul Barham, and Martin Abadi · 2013
Earlier work this paper cites.
End-to-end performance isolation through virtual datacenters
Sebastian Angel, Hitesh Ballani, Thomas Karagiannis, Greg O’Shea, and Eno Thereska · 2014
Earlier work this paper cites.
Hopper: Decentralized speculation-aware cluster scheduling at scale
Xiaoqi Ren, Ganesh Ananthanarayanan, Adam Wierman, and Minlan Yu · 2015
Earlier work this paper cites.
TensorFlow: A system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
vDNN: Virtualized deep neural networks for scalable, memory-efficient neural network design
Minsoo Rhu, Natalia Gimelshein, Jason Clemons, Arslan Zulfiqar, and Stephen W Keckler · 2016
Earlier work this paper cites.
Availability knob: Flexible user-defined availability in the cloud
Mohammad Shahrad and David Wentzlaff · 2016
Earlier work this paper cites.
Zorua: A holistic approach to resource virtualization in GPUs
Nandita Vijaykumar, Kevin Hsieh, Gennady Pekhimenko, Samira Khan, Ashish Shrestha, Saugata Ghose, Adwait Jog, Phillip B Gibbons, and Onur Mutlu · 2016
Earlier work this paper cites.
Clipper: A low-latency online prediction serving system
Daniel Crankshaw, Xin Wang, Guilio Zhou, Michael J Franklin, Joseph E Gonzalez, and Ion Stoica · 2017
Earlier work this paper cites.
Ultra-performance Pascal GPU and NVLink interconnect
Denis Foley and John Danskin · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Large batch training of convolutional networks
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Earlier work this paper cites.
Mask: Redesigning the GPU memory hierarchy to support multi-application concurrency
Rachata Ausavarungnirun, Vance Miller, Joshua Landgraf, Saugata Ghose, Jayneel Gandhi, Adwait Jog, Christopher J Rossbach, and Onur Mutlu · 2018
Earlier work this paper cites.
JAX: Composable transformations of Python+NumPy programs
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Earlier work this paper cites.
TVM: An automated end-to-end optimizing compiler for deep learning
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy · 2018
Earlier work this paper cites.
Shampoo: Preconditioned stochastic tensor optimization
Vineet Gupta, Tomer Koren, and Yoram Singer · 2018
Earlier work this paper cites.
Matrix capsules with EM routing
Geoffrey E Hinton, Sara Sabour, and Nicholas Frosst · 2018
Earlier work this paper cites.
Multi-tenant GPU clusters for deep learning workloads: Analysis and implications
Myeongjae Jeon, Shivaram Venkataraman, Junjie Qian, Amar Phanishayee, Wencong Xiao, and Fan Yang · 2018
Earlier work this paper cites.
Modeling task relationships in multi-task learning with multi-gate mixture-of-experts
Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H Chi · 2018
Cited alongside, same era.
Ray: A distributed framework for emerging AI applications
Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, Melih Elibol, Zongheng Yang, William Paul, Michael I. Jordan, and Ion Stoica · 2018
Cited alongside, same era.
Efficient neural architecture search via parameters sharing
Hieu Pham, Melody Guan, Barret Zoph, Quoc Le, and Jeff Dean · 2018
Cited alongside, same era.
Measuring the effects of data parallelism on neural network training
Christopher J Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E Dahl · 2018
Cited alongside, same era.
Mesh-TensorFlow: Deep learning for supercomputers
Noam Shazeer, Youlong Cheng, Niki Parmar, Dustin Tran, Ashish Vaswani, Penporn Koanantakool, Peter Hawkins, HyoukJoong Lee, Mingsheng Hong, Cliff Young, et al · 2018
Cited alongside, same era.
Neural message passing for multi-label classification
Jack Lanchantin, Arshdeep Sekhon, and Yanjun Qi · 2020
Later among the works it cites.
GShard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2020
Later among the works it cites.
Themis: Fair and efficient GPU cluster scheduling
Kshiteej Mahajan, Arjun Balasubramanian, Arjun Singhvi, Shivaram Venkataraman, Aditya Akella, Amar Phanishayee, and Shuchi Chawla · 2020
Later among the works it cites.
MLPerf: An industry standard benchmark suite for machine learning performance
Peter Mattson, Vijay Janapa Reddi, Christine Cheng, Cody Coleman, Greg Diamos, David Kanter, Paulius Micikevicius, David Patterson, Guenther Schmuelling, Hanlin Tang, Gu-Yeon Wei, and Carole-Jean Wu · 2020
Later among the works it cites.
Heterogeneity-aware cluster scheduling policies for deep learning workloads
Deepak Narayanan, Keshav Santhanam, Fiodar Kazhamiaka, Amar Phanishayee, and Matei Zaharia · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gandiva: Introspective cluster scheduling for deep learning
Wencong Xiao, Romil Bhardwaj, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra, Zhenhua Han, Pratyush Patel, Xuan Peng, Hanyu Zhao, Quanlu Zhang, et al · 2018
Cited alongside, same era.
Dynamic control flow in large-scale machine learning
Yuan Yu, Martin Abadi, Paul Barham, Eugene Brevdo, Mike Burrows, Andy Davis, Jeff Dean, Sanjay Ghemawat, Tim Harley, Peter Hawkins, Michael Isard, Manjunath Kudlur, Rajat Monga, Derek Murray, and Xiaoqiang Zheng · 2018
Cited alongside, same era.
TensorFlow Eager: A multi-stage, Python-embedded DSL for machine learning
Akshay Agrawal, Akshay Naresh Modi, Alexandre Passos, Allen Lavoie, Ashish Agarwal, Asim Shankar, Igor Ganichev, Josh Levenberg, Mingsheng Hong, Rajat Monga, et al · 2019
Cited alongside, same era.
Machine learning systems are stuck in a rut
Paul Barham and Michael Isard · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 2019
Cited alongside, same era.
GPipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al · 2019
Cited alongside, same era.
Later among the works it cites.
Deep learning training in Facebook data centers: Design of scale-up and scale-out systems
Maxim Naumov, John Kim, Dheevatsa Mudigere, Srinivas Sridharan, Xiaodong Wang, Whitney Zhao, Serhat Yilmaz, Changkyu Kim, Hector Yuen, Mustafa Ozdal, et al · 2020
Later among the works it cites.
DeepSpeed: System optimizations enable training deep learning models with over 100 billion parameters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He · 2020
Later among the works it cites.
Antman: Dynamic scaling on GPU clusters for deep learning
Wencong Xiao, Shiru Ren, Yong Li, Yang Zhang, Pengyang Hou, Zhi Li, Yihui Feng, Wei Lin, and Yangqing Jia · 2020
Later among the works it cites.
AvA: Accelerated virtualization of accelerators
Hangchen Yu, Arthur Michener Peters, Amogh Akshintala, and Christopher J Rossbach · 2020
Later among the works it cites.
Fine-grained GPU sharing primitives for deep learning applications
Peifeng Yu and Mosharaf Chowdhury · 2020
Later among the works it cites.
Scalable second order optimization for deep learning
Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, and Yoram Singer · 2021
Later among the works it cites.
On the opportunities and risks of foundation models
Rishi Bommasani and Drew A. Hudson et. al · 2021
Later among the works it cites.
Introducing Pathways: A next-generation AI architecture
Jeff Dean · 2021
Later among the works it cites.
Switch Transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Later among the works it cites.
Cloud TPU
Google · 2021
Later among the works it cites.
MLIR: Scaling compiler infrastructure for domain specific computation
Chris Lattner, Mehdi Amini, Uday Bondhugula, Albert Cohen, Andy Davis, Jacques Pienaar, River Riddle, Tatiana Shpeisman, Nicolas Vasilache, and Oleksandr Zinenko · 2021
Later among the works it cites.
Zico: Efficient GPU memory sharing for concurrent DNN training
Gangmuk Lim, Jeongseob Ahn, Wencong Xiao, Youngjin Kwon, and Myeongjae Jeon · 2021
Later among the works it cites.
Efficient large-scale language model training on GPU clusters using Megatron-LM
Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Amar Phanishayee, and Matei Zaharia · 2021
Later among the works it cites.
NVIDIA GPUDirect technology
NVIDIA · 2021
Later among the works it cites.
ZeRO-Infinity: Breaking the GPU memory wall for extreme scale deep learning
Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, and Yuxiong He · 2021
Later among the works it cites.
TensorFlow Datasets: A collection of ready-to-use datasets
TensorFlow · 2021
Later among the works it cites.
Wavelet: Efficient DNN training with Tick-Tock scheduling
Guanhua Wang, Kehan Wang, Kenan Jiang, Xiangjun Li, and Ion Stoica · 2021
Later among the works it cites.
Pipemare: Asynchronous pipeline parallel DNN training
Bowen Yang, Jian Zhang, Jonathan Li, Christopher Ré, Christopher Aberger, and Christopher De Sa · 2021
Later among the works it cites.
Share or not? learning to schedule language-specific capacity for multilingual translation
Biao Zhang, Ankur Bapna, Rico Sennrich, and Orhan Firat · 2021
Later among the works it cites.
MLaaS in the wild: Workload analysis and scheduling in large-scale heterogeneous GPU clusters
Qizhen Weng, Wencong Xiao, Yinghao Yu, Wei Wang, Cheng Wang, Jian He, Yong Li, Liping Zhang, Wei Lin, and Yu Ding · 2022
Closest in time.
vPipe: A virtualized acceleration system for achieving efficient and scalable pipeline parallel DNN training
Shixiong Zhao, Fanxin Li, Xusheng Chen, Xiuxian Guan, Jianyu Jiang, Dong Huang, Yuhao Qing, Sen Wang, Peng Wang, Gong Zhang, Cheng Li, Ping Luo, and Heming Cui · 2022
Closest in time.