Fetching the paper…
Reading the bibliography…
Checkpointing enables the training of deep learning models under restricted memory budgets by freeing intermediate activations from memory and recomputing them on demand.
Low-memory neural network training: A technical report
Nimit Sharad Sohoni, Christopher Richard Aberger, Megan Leszczynski, Jian Zhang, and Christopher Ré · 1904
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library, 2019
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 1912
Earlier work this paper cites.
Achieving logarithmic growth of temporal and spatial complexity in reverse automatic differentiation
Andreas Griewank · 1994
Earlier work this paper cites.
Space-efficient inference in dynamic probabilistic networks
John Binder, Kevin Murphy, and Stuart Russell · 1997
Earlier work this paper cites.
Treeverse: An implementation of checkpointing for the reverse or adjoint mode of computational differentiation
Andreas Griewank and Andrea Walther · 1997
Earlier work this paper cites.
Algorithm 799: Revolve: An implementation of checkpoint for the reverse or adjoint mode of computational differentiation
Andreas Griewank and Andrea Walther · 2000
Earlier work this paper cites.
The tapenade automatic differentiation tool: Principles, model, and specification
Laurent Hascoet and Valérie Pascual · 2013
Earlier work this paper cites.
Automatic differentiation in machine learning: a survey
Atilim Gunes Baydin, Barak A. Pearlmutter, Alexey Andreyevich Radul, and Jeffrey Mark Siskind · 2015
Earlier work this paper cites.
Improved semantic representations from tree-structured long short-term memory networks
Kai Sheng Tai, Richard Socher, and Christopher D. Manning · 2015
Cited alongside, same era.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin · 2016
Cited alongside, same era.
Memory-efficient backpropagation through time
Audrunas Gruslys, Rémi Munos, Ivo Danihelka, Marc Lanctot, and Alex Graves · 2016
Cited alongside, same era.
The reversible residual network: Backpropagation without storing activations
Aidan N Gomez, Mengye Ren, Raquel Urtasun, and Roger B Grosse · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
Optimal memory-aware backpropagation of deep join networks
Olivier Beaumont, Julien Herrmann, Guillaume Pallez, and Alena Shilova · 2019
Later among the works it cites.
Efficient rematerialization for deep networks
Ravi Kumar, Manish Purohit, Zoya Svitkina, Erik Vee, and Joshua Wang · 2019
Later among the works it cites.
A graph theoretic framework of recomputation algorithms for memory-efficient backpropagation
Mitsuru Kusumoto, Takuya Inoue, Gentaro Watanabe, Takuya Akiba, and Masanori Koyama · 2019
Later among the works it cites.
Swapadvisor: Pushing deep learning beyond the gpu memory limit via smart swapping
Chien-Chin Huang, Gu Jin, and Jinyang Li · 2020
Closest in time.
Checkmate: Breaking the memory wall with optimal tensor rematerialization
Paras Jain, Ajay Jain, Aniruddha Nrusimha, Amir Gholami, Pieter Abbeel, Joseph Gonzalez, Kurt Keutzer, and Ion Stoica · 2020
Closest in time.
Reformer: The efficient transformer
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Glow: Graph lowering compiler techniques for neural networks
Nadav Rotem, Jordan Fix, Saleem Abdulrasool, Summer Deng, Roman Dzhabarov, James Hegeman, Roman Levenstein, Bert Maher, Satish Nadathur, Jakob Olesen, Jongsoo Park, Artem Rakhov, and Misha Smelyanskiy · 2018
Cited alongside, same era.
Divide-and-conquer checkpointing for arbitrary programs with no user annotation
Jeffrey Mark Siskind and Barak A. Pearlmutter · 2018
Cited alongside, same era.
Superneurons
Linnan Wang, Jinmian Ye, Yiyang Zhao, Wei Wu, Ang Li, Shuaiwen Leon Song, Zenglin Xu, and Tim Kraska · 2018
Cited alongside, same era.
Optimal checkpointing for heterogeneous chains: how to train deep neural networks with limited memory
Olivier Beaumont, Lionel Eyraud-Dubois, Julien Herrmann, Alexis Joly, and Alena Shilova
Cited in the paper.
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya · 2020
Closest in time.
Capuchin: Tensor-based gpu memory management for deep learning
Xuan Peng, Xuanhua Shi, Hulin Dai, Hai Jin, Weiliang Ma, Qian Xiong, Fan Yang, and Xuehai Qian · 2020
Closest in time.