Fetching the paper…
Reading the bibliography…
How does one compile derivatives of tensor programs, such that the resulting code is purely functional (hence easier to optimize and parallelize) and provably efficient relative to the original program? We show that naively differentiating tensor code---as done in popular systems like Tensorflow and PyTorch---can cause asymptotic slowdowns in pathological cases, violating the Cheap Gradients Principle.
A programming language. In Proceedings of the spring joint computer conference . 345–351
Kenneth E Iverson. 1962 · 1962
Earlier work this paper cites.
A Simple Automatic Derivative Evaluation Program
R. E. Wengert. 1964 · 1964
Earlier work this paper cites.
The representation of the cumulative rounding error of an algorithm as a Taylor expansion of the local rounding errors
Seppo Linnainmaa. 1970 · 1970
Earlier work this paper cites.
Computational graphs and rounding error
Friedrich L Bauer. 1974 · 1974
Earlier work this paper cites.
Lambda Lifting: Transforming Programs to Recursive Equations. Springer-Verlag, 190–203
Thomas Johnsson. 1985 · 1985
Earlier work this paper cites.
Concrete Mathematics: A Foundation for Computer Science
Ronald L. Graham, Donald E. Knuth, and Oren Patashnik. 1989 · 1989
Earlier work this paper cites.
On the Calculation of Jacobian Matrices by the Markowitz Rule
Andreas Griewank and Shawn Reese. 1991 · 1991
Earlier work this paper cites.
ADIFOR: Automatic Differentiation in a Source Translator Environment. In International Symposium on Symbolic and Algebraic Computation . 294–302
Christian Bischof, Alan Carle, George Corliss, and Andreas Griewank. 1992 · 1992
Earlier work this paper cites.
Efficient calculation of Jacobian matrices by optimized application of the chain rule to computational graphs
Uwe Naumann. 1999 · 1999
Earlier work this paper cites.
ADrien: an implementation of automatic differentiation in Maple. In International Symposium on Symbolic and Algebraic Computation . 221–228
Dominique Villard and Michael B Monagan. 1999 · 1999
Earlier work this paper cites.
Efficient symbolic differentiation for graphics applications
Brian K Guenter. 2007 · 2007
Earlier work this paper cites.
A practical automatic polyhedral parallelizer and locality optimizer. In SIGPLAN Not. , Vol. 43. ACM, 101–113
Uday Bondhugula, Albert Hartono, Jagannathan Ramanujam, and Ponnuswamy Sadayappan. 2008 · 2008
Earlier work this paper cites.
Evaluating Derivatives: Principles and Techniques of Algorithmic Differentiation (second ed.)
Andreas Griewank and Andrea Walther. 2008 · 2008
Earlier work this paper cites.
Optimal Jacobian accumulation is NP-complete
Uwe Naumann. 2008 · 2008
Earlier work this paper cites.
Reverse-mode AD in a Functional Framework: Lambda the Ultimate Backpropagator
Barak A. Pearlmutter and Jeffrey Mark Siskind. 2008 · 2008
Earlier work this paper cites.
OpenAD/F: A modular open-source tool for automatic differentiation of Fortran codes
Jean Utke, Uwe Naumann, Mike Fagan, Nathan Tallent, Michelle Strout, Patrick Heimbach, Chris Hill, and Carl Wunsch. 2008 · 2008
Earlier work this paper cites.
Beautiful differentiation
Conal M Elliott. 2009 · 2009
Earlier work this paper cites.
Theano: a CPU and GPU Math Expression Compiler. In Python for Scientific Computing Conference (SciPy)
James Bergstra, Olivier Breuleux, Frédéric Bastien, Pascal Lamblin, Razvan Pascanu, Guillaume Desjardins, Joseph Turian, David Warde-Farley, and Yoshua Bengio. 2010 · 2010
Earlier work this paper cites.
Who Invented the Reverse Mode of Differentiation?
Andreas Griewank. 2012 · 2012
Cited alongside, same era.
Polly — performing polyhedral optimizations on a low-level intermediate representation
Tobias Grosser, Armin Groesslinger, and Christian Lengauer. 2012 · 2012
Cited alongside, same era.
Decoupling Algorithms from Schedules for Easy Optimization of Image Processing Pipelines
Jonathan Ragan-Kelley, Andrew Adams, Sylvain Paris, Marc Levoy, Saman Amarasinghe, and Frédo Durand. 2012 · 2012
Cited alongside, same era.
The Tapenade Automatic Differentiation Tool: Principles, Model, and Specification
Laurent Hascoet and Valérie Pascual. 2013 · 2013
Cited alongside, same era.
Halide: A Language and Compiler for Optimizing Parallelism, Locality, and Recomputation in Image Processing Pipelines
Jonathan Ragan-Kelley, Connelly Barnes, Andrew Adams, Sylvain Paris, Frédo Durand, and Saman Amarasinghe. 2013 · 2013
Cited alongside, same era.
Relay: A new ir for machine learning frameworks. In Proceedings of the 2nd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages . 58–68
Jared Roesch, Steven Lyubomirsky, Logan Weber, Josh Pollock, Marisa Kirisame, Tianqi Chen, and Zachary Tatlock. 2018 · 2018
Later among the works it cites.
Tensor comprehensions: Framework-agnostic high-performance machine learning abstractions
Nicolas Vasilache, Oleksandr Zinenko, Theodoros Theodoridis, Priya Goyal, Zachary DeVito, William S Moses, Sven Verdoolaege, Andrew Adams, and Albert Cohen. 2018 · 2018
Later among the works it cites.
A Simple Differentiable Programming Language
Martín Abadi and Gordon D. Plotkin. 2019 · 2019
Later among the works it cites.
Machine learning systems are stuck in a rut. In Proceedings of the Workshop on Hot Topics in Operating Systems . 177–183
Paul Barham and Michael Isard. 2019 · 2019
Later among the works it cites.
Backpropagation in the simply typed Lambda-calculus with linear negation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2015 · 2015
Cited alongside, same era.
CVXPY: A Python-embedded modeling language for convex optimization
Steven Diamond and Stephen Boyd. 2016 · 2016
Cited alongside, same era.
Proximal: Efficient image optimization using proximal algorithms
Felix Heide, Steven Diamond, Matthias Nießner, Jonathan Ragan-Kelley, Wolfgang Heidrich, and Gordon Wetzstein. 2016 · 2016
Cited alongside, same era.
Edge Pushing is Equivalent to Vertex Elimination for Computing Hessians. In Workshop on Combinatorial Scientific Computing . SIAM, 102–111
Mu Wang, Alex Pothen, and Paul Hovland. 2016 · 2016
Cited alongside, same era.
Opt: A Domain Specific Language for Non-Linear Least Squares Optimization in Graphics and Imaging
Zachary Devito, Michael Mara, Michael Zollhöfer, Gilbert Bernstein, Jonathan Ragan-Kelley, Christian Theobalt, Pat Hanrahan, Matthew Fisher, and Matthias Niessner. 2017 · 2017
Cited alongside, same era.
The tensor algebra compiler
Fredrik Kjolstad, Shoaib Kamil, Stephen Chou, David Lugato, and Saman Amarasinghe. 2017 · 2017
Cited alongside, same era.
A rewriting system for convex optimization problems
Akshay Agrawal, Robin Verschueren, Steven Diamond, and Stephen Boyd. 2018 · 2018
Cited alongside, same era.
Aloïs Brunel, Damiano Mazza, and Michele Pagani. 2019a · 2019
Later among the works it cites.
Backpropagation in the Simply Typed Lambda-Calculus with Linear Negation
Aloïs Brunel, Damiano Mazza, and Michele Pagani. 2019b · 2019
Later among the works it cites.
Taichi: a language for high-performance computation on spatially sparse data structures
Yuanming Hu, Tzu-Mao Li, Luke Anderson, Jonathan Ragan-Kelley, and Frédo Durand. 2019 · 2019
Later among the works it cites.
Towards Polyhedral Automatic Differentiation. In NeurIPS Workshop – Program Transformations for Machine Learning
Jan Hückelheim and Navjot Kukreja. 2019 · 2019
Later among the works it cites.
Automatic differentiation for adjoint stencil loops. In Proceedings of the International Conference on Parallel Processing . 1–10
Jan Hückelheim, Navjot Kukreja, Sri Hari Krishna Narayanan, Fabio Luporini, Gerard Gorman, and Paul Hovland. 2019 · 2019
Later among the works it cites.
TASO: Optimizing Deep Learning Computation with Automatic Generation of Graph Substitutions. In Proceedings of the ACM Symposium on Operating Systems Principles (SOSP) . ACM, 47–62
Zhihao Jia, Oded Padon, James Thomas, Todd Warszawski, Matei Zaharia, and Alex Aiken. 2019 · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems . 8026–8037
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
Efficient Differentiable Programming in a Functional Array-Processing Language
Amir Shaikhha, Andrew Fitzgibbon, Dimitrios Vytiniotis, and Simon Peyton Jones. 2019 · 2019
Later among the works it cites.
Demystifying differentiable programming: Shift/reset the penultimate backpropagator
Fei Wang, Daniel Zheng, James Decker, Xilun Wu, Grégory M Essertel, and Tiark Rompf. 2019 · 2019
Later among the works it cites.
DiffTaichi: Differentiable Programming for Physical Simulation
Yuanming Hu, Luke Anderson, Tzu-Mao Li, Qi Sun, Nathan Carr, Jonathan Ragan-Kelley, and Frédo Durand. 2020 · 2020
Closest in time.
How Are Convolutions Actually Performed Under the Hood?
Anirudh Shenoy. 2019 · 2020
Closest in time.
Benjamin Sherman, Jesse Michel, and Michael Carbin. 2020 · 2020
Closest in time.
XLA – TensorFlow compiled
The XLA Team. 2017 · 2020
Closest in time.