Fetching the paper…
Reading the bibliography…
This paper presents Tofu, a system that partitions very large DNN models across multiple GPU devices to reduce per-GPU memory footprint.
A cellular computer to implement the Kalman Filter Algorithm
L. E. Cannon · 1969
Earlier work this paper cites.
Graph Theory with Applications
J.A Bondy and U.S.R. Murty · 1976
Earlier work this paper cites.
A methodology for parallelizing programs for multicomputers and complex memory multiprocessors
J Ramanujam and P Sadayappan · 1989
Earlier work this paper cites.
Partitioning and labeling of index sets in do loops with constant dependence vectors
ERIKH D’HOLLANDER · 1989
Earlier work this paper cites.
Index domain alignment: Minimizing cost of cross-referencing between distributed arrays
Jingke Li and Marina Chen · 1990
Earlier work this paper cites.
LAPACK: A portable linear algebra library for high-performance computers
Edward Anderson, Zhaojun Bai, J Dongarra, A Greenbaum, A McKenney, Jeremy Du Croz, S Hammerling, J Demmel, C Bischof, and Danny Sorensen · 1990
Earlier work this paper cites.
Compiler techniques for data partitioning of sequentially iterated parallel loops
David E Hudak and Santosh G Abraham · 1990
Earlier work this paper cites.
Data optimization: Allocation of arrays to reduce communication on simd machines
Kathleen Knobe, Joan D Lukas, and Guy L Steele Jr · 1990
Earlier work this paper cites.
The data alignment phase in compiling programs for distributed-memory machines
Jingke Li and Marina Chen · 1991
Earlier work this paper cites.
Compile-time techniques for data distribution in distributed memory machines
J Ramanujam and P Sadayappan · 1991
Earlier work this paper cites.
Scalapack: A scalable linear algebra library for distributed memory concurrent computers
Jaeyoung Choi, Jack J Dongarra, Roldan Pozo, and David W Walker · 1992
Earlier work this paper cites.
Reduction of cache coherence overhead by compiler data layout and loop transformation
Y-J Ju and H Dietz · 1992
Earlier work this paper cites.
Np-completeness of dynamic remapping
Ulrich Kremer · 1993
Earlier work this paper cites.
Communication-free hyperplane partitioning of nested loops
Chua-Huang Huang and Ponnuswamy Sadayappan · 1993
Earlier work this paper cites.
ZPL: An array sublanguage
Calvin Lin and Lawrence Snyder · 1994
Earlier work this paper cites.
Summa: Scalable universal matrix multiplication algorithm
Robert A. van de Geijn and Jerrell Watts · 1995
Earlier work this paper cites.
Automatic alignment of array data and processes to reduce communication time on DMPPs
Michael Philippsen · 1995
Earlier work this paper cites.
Solving alignment using elementary linear algebra
David Bau, Induprakas Kodukula, Vladimir Kotlyar, Keshav Pingali, and Paul Stodghill · 1995
Earlier work this paper cites.
Global arrays: A nonuniform memory access programming model for high-performance computers
Jaroslaw Nieplocha, Robert J Harrison, and Richard J Littlefield · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Efficient management of parallelism in object oriented numerical software libraries
Satish Balay, William D. Gropp, Lois Curfman McInnes, and Barry F. Smith · 1997
Earlier work this paper cites.
Automatic data layout for distributed-memory machines
Ken Kennedy and Ulrich Kremer · 1998
Earlier work this paper cites.
Automatic array alignment in parallel matlab scripts
Igor Z Milosavljevic and Marwan A Jabri · 1999
Earlier work this paper cites.
Symbolic bounds analysis of pointers, array indices, and accessed memory regions
Radu Rugina and Martin Rinard · 2000
Earlier work this paper cites.
Tensor contraction engine: Abstraction and automated parallel implementation of configuration-interaction, coupled-cluster, and many-body perturbation theories
So Hirata · 2003
Earlier work this paper cites.
Mapreduce: Simplified data processing on large clusters
Jeff Dean and Sanjay Ghemawat · 2004
Earlier work this paper cites.
UPC language specifications, v1.2
UPC Consortium · 2005
Earlier work this paper cites.
Parallel programmability and the chapel language
B.L. Chamberlain, D. Callahan, and H.P. Zima · 2007
Earlier work this paper cites.
Data layout transformation for enhancing data locality on nuca chip multiprocessors
Qingda Lu, Christophe Alias, Uday Bondhugula, Thomas Henretty, Sriram Krishnamoorthy, Jagannathan Ramanujam, Atanas Rountev, Ponnuswamy Sadayappan, Yongjian Chen, Haibo Lin, et al · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Cited alongside, same era.
Scihadoop: array-based query processing in hadoop
Joe B. Buck, Noah Watkins, Jeff LeFevre, Kleoni Ioannidou, Carlos Maltzahn, Neoklis Polyzotis, and Scott Brandt · 2011
Cited alongside, same era.
Large-scale parallel statistical forecasting computations in r
Murray Stokely, Farzan Rohani, and Eric Tassone · 2011
Cited alongside, same era.
Training deep and recurrent networks with hessian-free optimization
James Martens and Ilya Sutskever · 2012
Cited alongside, same era.
Large scale distributed deep networks
Jeffrey Dean, Greg S. Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Quoc V. Le, Mark Z. Mao, Marc’Aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, and Andrew Y. Ng · 2012
Cited alongside, same era.
MadLINQ: large-scale distributed matrix computation for the cloud
Sergey Zagoruyko and Nikos Komodakis · 2016
Later among the works it cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, and Mohammad Norouzi · 2016
Later among the works it cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Later among the works it cites.
Memory-efficient backpropagation through time
Audrunas Gruslys, Rémi Munos, Ivo Danihelka, Marc Lanctot, and Alex Graves · 2016
Later among the works it cites.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhengping Qian, Xiuwei Chen, Nanxi Kang, Mingcheng Chen, Yuan Yu, Thomas Moscibroda, and Zheng Zhang · 2012
Cited alongside, same era.
The gauge domain: scalable analysis of linear inequality invariants
Arnaud J Venet · 2012
Cited alongside, same era.
Powergraph: Distributed graph-parallel computation on natural graphs
Joseph E. Gonzalez, Yucheng Low, Haijie Gu, Danny Bickson, and Carlos Guestrin · 2012
Cited alongside, same era.
Deep learning with COTS HPC systems
Adam Coates, Brody Huval, Tao Wang, David Wu, Bryan Catanzaro, and Ng Andrew · 2013
Cited alongside, same era.
Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines
Jonathan Ragan-Kelley, Connelly Barnes, Andrew Adams, Sylvain Paris, Frédo Durand, and Saman Amarasinghe · 2013
Cited alongside, same era.
Elemental: A new framework for distributed memory dense matrix computations
Jack Poulson, Bryan Marker, Robert A. van de Geijn, Jeff R. Hammond, and Nichols A. Romero · 2013
Cited alongside, same era.
Presto: distributed machine learning and graph processing with sparse matrices
Shivaram Venkataraman, Erik Bodzsar, Indrajit Roy, Alvin AuYoung, and Robert S. Schreiber · 2013
Cited alongside, same era.
Xuan Yang, Jing Pu, Blaine Burton Rister, Nikhil Bhagdikar, Stephen Richardson, Shahar Kvatinsky, Jonathan Ragan-Kelley, Ardavan Pedram, and Mark Horowitz · 2016
Later among the works it cites.
Neurocube: A programmable digital neuromorphic architecture with high-density 3d memory
Duckhwan Kim, Jaeha Kung, Sek Chai, Sudhakar Yalamanchili, and Saibal Mukhopadhyay · 2016
Later among the works it cites.
Tensorflow: A system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2016
Later among the works it cites.
Exploring the limits of language modeling
Rafal Józefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu · 2016
Later among the works it cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Later among the works it cites.
vdnn: Virtualized deep neural networks for scalable, memory-efficient neural network design
Minsoo Rhu, Natalia Gimelshein, Jason Clemons, Arslan Zulfiqar, and Stephen W Keckler · 2016
Later among the works it cites.
Geeps: Scalable deep learning on distributed gpus with a gpu-specialized parameter server
Henggang Cui, Hao Zhang, Gregory R. Ganger, Phillip B. Gibbons, and Eric P. Xing · 2016
Later among the works it cites.
Strads: A distributed framework for scheduled model parallel machine learning
Jin Kyu Kim, Qirong Ho, Seunghak Lee, Xun Zheng, Wei Dai, Garth Gibson, and Eric Xing · 2016
Later among the works it cites.
Binarized neural networks
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio · 2016
Later among the works it cites.
Measuring and optimizing distributed array programs
Mingxing Zhang, Yongwei Wu, Kang Chen, Teng Ma, and Weimin Zheng · 2016
Later among the works it cites.
Have abstraction and eat performance, too: Optimized heterogeneous computing with parallel patterns
Kevin J. Brown, HyoukJoong Lee, Tiark Rompf, Arvind K. Sujeeth, Christopher De Sa, Christopher Aberger, and Kunle Olukotun · 2016
Later among the works it cites.
Distributed halide
Tyler Denniston, Shoaib Kamil, and Saman Amarasinghe · 2016
Later among the works it cites.
Training deeper models by gpu memory optimization on tensorflow
Chen Meng, Minmin Sun, Jun Yang, Minghui Qiu, and Yang Gu · 2017
Later among the works it cites.
Tetris: Scalable and efficient neural network acceleration with 3d memory
Mingyu Gao, Jing Pu, Xuan Yang, Mark Horowitz, and Christos Kozyrakis · 2017
Later among the works it cites.
Device placement optimization with reinforcement learning
Azalia Mirhoseini, Hieu Pham, Quoc V Le, Benoit Steiner, Rasmus Larsen, Yuefeng Zhou, Naveen Kumar, Mohammad Norouzi, Samy Bengio, and Jeff Dean · 2017
Later among the works it cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Later among the works it cites.
Exploring hidden dimensions in parallelizing convolutional neural networks
Zhihao Jia, Sina Lin, Charles R. Qi, and Alex Aiken · 2018
Closest in time.
Beyond data and model parallelism for deep neural networks
Zhihao Jia, Matei Zaharia, and Alex Aiken · 2018
Closest in time.
A hierarchical model for device placement
Azalia Mirhoseini, Anna Goldie, Hieu Pham, Benoit Steiner, Quoc V. Le, and Jeff Dean · 2018
Closest in time.
TVM: An automated end-to-end optimizing compiler for deep learning
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy · 2018
Closest in time.
Tensor comprehensions: Framework-agnostic high-performance machine learning abstractions
Nicolas Vasilache, Oleksandr Zinenko, Theodoros Theodoridis, Priya Goyal, Zachary DeVito, William S. Moses, Sven Verdoolaege, Andrew Adams, and Albert Cohen · 2018
Closest in time.
Supporting very large models using automatic dataflow graph partitioning
Minjie Wang, Chien-chin Huang, and Jinyang Li · 2018
Closest in time.
Profile-guided memory optimization for deep neural networks
Taro Sekiyama, Takashi Imamichi, Haruki Imai, and Rudy Raymond · 2018
Closest in time.