Fetching the paper…
Reading the bibliography…
We formalize the problem of trading-off DNN training time and memory requirements as the tensor rematerialization optimization problem, a generalization of prior checkpointing strategies.
Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 1901
Earlier work this paper cites.
Generating Long Sequences with Sparse Transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 1904
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 1907
Earlier work this paper cites.
On Compiling Algorithms for Arithmetic Expressions
Nakata, I · 1967
Earlier work this paper cites.
Complete Register Allocation Problems
Sethi, R · 1973
Earlier work this paper cites.
Register allocation via coloring
Chaitin, G. J., Auslander, M. A., Chandra, A. K., Cocke, J., Hopkins, M. E., and Markstein, P. W · 1981
Earlier work this paper cites.
A new polynomial-time algorithm for linear programming
Karmarkar, N · 1984
Earlier work this paper cites.
Global Value Numbers and Redundant Computations
Rosen, B. K., Wegman, M. N., and Zadeck, F. K · 1988
Earlier work this paper cites.
Efficiently Computing Static Single Assignment Form and the Control Dependence Graph
Cytron, R., Ferrante, J., Rosen, B. K., Wegman, M. N., and Zadeck, F. K · 1991
Earlier work this paper cites.
Rematerialization
Briggs, P., Cooper, K. D., and Torczon, L · 1992
Earlier work this paper cites.
Interior-point polynomial algorithms in convex programming , volume 13
Nesterov, Y. and Nemirovskii, A · 1994
Earlier work this paper cites.
On the approximation of maximum satisfiability
Yannakakis, M · 1994
Earlier work this paper cites.
Optimal and Near-optimal Global Register Allocation Using 0–1 Integer Programming
Goodwin, D. W. and Wilken, K. D · 1996
Earlier work this paper cites.
Algorithm 799: revolve: an implementation of checkpointing for the reverse or adjoint mode of computational differentiation
Griewank, A. and Walther, A · 2000
Earlier work this paper cites.
LLVM: An Infrastructure for Multi-Stage Optimization
Lattner, C · 2002
Earlier work this paper cites.
Register Rematerialization in GCC
Punjani, M · 2004
Earlier work this paper cites.
A Global Progressive Register Allocator
Koes, D. R. and Goldstein, S. C · 2006
Earlier work this paper cites.
Graph Algorithms: Applications, 2008
Holder, L · 2008
Earlier work this paper cites.
Local Memory and Register Spilling, 2011
Micikevicius, P · 2011
Earlier work this paper cites.
Register Allocation in LLVM 3.0, November 2011
Olesen, J. S · 2011
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Cited alongside, same era.
Very Deep Convolutional Networks for Large-Scale Image Recognition
Simonyan, K. and Zisserman, A · 2014
Cited alongside, same era.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
Fully convolutional networks for semantic segmentation
Long, J., Shelhamer, E., and Darrell, T · 2015
Cited alongside, same era.
U-Net: Convolutional Networks for Biomedical Image Segmentation
Ronneberger, O., Fischer, P., and Brox, T · 2015
Cited alongside, same era.
Going deeper with convolutions
Szegedy, C., Wei Liu, Yangqing Jia, Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Image super-resolution via deep recursive residual network
Tai, Y., Yang, J., and Liu, X · 2017
Later among the works it cites.
Attention is All you Need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Later among the works it cites.
Aggregated residual transformations for deep neural networks
Xie, S., Girshick, R., Dollár, P., Tu, Z., and He, K · 2017
Later among the works it cites.
Large scale GAN training for high fidelity natural image synthesis
Brock, A., Donahue, J., and Simonyan, K · 2018
Later among the works it cites.
In-place Activated BatchNorm for Memory-Optimized Training of DNNs
Bulo, S. R., Porzi, L., and Kontschieder, P · 2018
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mane, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viegas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2016
Cited alongside, same era.
An Analysis of Deep Neural Network Models for Practical Applications
Canziani, A., Paszke, A., and Culurciello, E · 2016
Cited alongside, same era.
Image super-resolution using deep convolutional networks
Dong, C., Loy, C. C., He, K., and Tang, X · 2016
Cited alongside, same era.
Memory-efficient Backpropagation Through Time
Gruslys, A., Munos, R., Danihelka, I., Lanctot, M., and Graves, A · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Accurate image super-resolution using very deep convolutional networks
Kim, J., Lee, J. K., and Lee, K. M · 2016
Cited alongside, same era.
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Later among the works it cites.
Cutting Down Training Memory by Re-fowarding
Feng, J. and Huang, D · 2018
Later among the works it cites.
Integrated model, batch, and domain parallelism in training neural networks
Gholami, A., Azad, A., Jin, P., Keutzer, K., and Buluc, A · 2018
Later among the works it cites.
Faster Neural Networks Straight from JPEG
Gueguen, L., Sergeev, A., Kadlec, B., Liu, R., and Yosinski, J · 2018
Later among the works it cites.
Gist: Efficient Data Encoding for Deep Neural Network Training
Jain, A., Phanishayee, A., Mars, J., Tang, L., and Pekhimenko, G · 2018
Later among the works it cites.
Combinatorial Register Allocation and Instruction Scheduling
Lozano, R. C., Carlsson, M., Blindell, G. H., and Schulte, C · 2018
Later among the works it cites.
Divide-and-conquer checkpointing for arbitrary programs with no user annotation
Siskind, J. M. and Pearlmutter, B. A · 2018
Later among the works it cites.
Group Normalization
Wu, Y. and He, K · 2018
Later among the works it cites.
HDNET: Exploiting HD Maps for 3D Object Detection
Yang, B., Liang, M., and Urtasun, R · 2018
Later among the works it cites.
Optimal memory-aware backpropagation of deep join networks
Beaumont, O., Herrmann, J., Pallez, G., and Shilova, A · 2019
Closest in time.
COIN-OR Branch-and-Cut solver, June 2019
Forrest, J. J., Vigerske, S., Ralphs, T., Santos, H. G., Hafer, L., Kristjansson, B., Fasano, J., Straver, E., Lubin, M., rlougee, jpgoncal1, Gassmann, H. I., and Saltzman, M · 2019
Closest in time.
The ooo vliw jit compiler for gpu inference
Jain, P., Mo, X., Jain, A., Tumanov, A., Gonzalez, J. E., and Stoica, I · 2019
Closest in time.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Closest in time.
Astra: Exploiting Predictability to Optimize Deep Learning
Sivathanu, M., Chugh, T., Singapuram, S. S., and Zhou, L · 2019
Closest in time.