Fetching the paper…
Reading the bibliography…
Derivatives, mostly in the form of gradients and Hessians, are ubiquitous in machine learning.
Analytical differentiation on a digital computer
John F. Nolan · 1953
Earlier work this paper cites.
Programs for automatic differentiation for the machine BESM (in Russian)
L. M. Beda, L. N. Korolev, N. V. Sukkikh, and T. S. Frolova · 1959
Earlier work this paper cites.
L. S. Pontryagin’s maximum principle in the theory of optimum systems—Part II
L. I. Rozonoer · 1959
Earlier work this paper cites.
The theory of optimal processes I: The maximum principle
V. G. Boltyanskii, R. V. Gamkrelidze, and L. S. Pontryagin · 1960
Earlier work this paper cites.
A steepest ascent method for solving optimum programming problems
A. E. Bryson and W. F. Denham · 1962
Earlier work this paper cites.
A simple automatic derivative evaluation program
Robert E. Wengert · 1964
Earlier work this paper cites.
A SLANG simulation of an initially strong shock wave downstream of an infinite area change
D. S. Adamson and C. W. Winant · 1969
Earlier work this paper cites.
Applied Optimal Control: Optimization, Estimation, and Control
Arthur E. Bryson and Yu-Chi Ho · 1969
Earlier work this paper cites.
The representation of the cumulative rounding error of an algorithm as a taylor expansion of the local rounding errors
Seppo Linnainmaa · 1970
Earlier work this paper cites.
Differential Dynamic Programming
David Q. Mayne and David H. Jacobson · 1970
Earlier work this paper cites.
Computing derivatives using W-arithmetic and U-arithmetic
C. L. Lawson · 1971
Earlier work this paper cites.
Computational graphs and rounding error
Friedrich L. Bauer · 1974
Earlier work this paper cites.
Beyond Regression: New Tools for Prediction and Analysis in the Behavioral Sciences
Paul J. Werbos · 1974
Earlier work this paper cites.
Taylor expansion of the accumulated rounding error
Seppo Linnainmaa · 1976
Earlier work this paper cites.
A new two-constant equation of state
Ding-Yu Peng and Donald B. Robinson · 1976
Earlier work this paper cites.
Understanding image intensities
Berthold K. P. Horn · 1977
Earlier work this paper cites.
Compiling Fast Partial Derivatives of Functions Given by Algorithms
Bert Speelpenning · 1980
Earlier work this paper cites.
Numerical differentiation of analytic functions
Bengt Fornberg · 1981
Earlier work this paper cites.
Learning-logic: Casting the cortex of the human brain in silicon
David B. Parker · 1985
Earlier work this paper cites.
Learning representations by back-propagating errors
David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams · 1986
Earlier work this paper cites.
Hybrid Monte Carlo
Simon Duane, Anthony D. Kennedy, Brian J. Pendleton, and Duncan Roweth · 1987
Earlier work this paper cites.
Automatic differentiation in PROSE
F. W. Pfeiffer · 1987
Earlier work this paper cites.
SN: A simulator for connectionist models
Léon Bottou and Yann LeCun · 1988
Earlier work this paper cites.
Application of differentiation arithmetic , volume 19 of Perspectives in Computing , pages 127–48
George F. Corliss · 1988
Earlier work this paper cites.
GRESS version 1.0 user’s manual
Jim E. Horwedel, Brian A. Worley, E. M. Oblow, and F. G. Pin · 1988
Earlier work this paper cites.
Runtime tags aren’t necessary
Andrew W Appel · 1989
Earlier work this paper cites.
On automatic differentiation
Andreas Griewank · 1989
Earlier work this paper cites.
Theory of the backpropagation neural network
Robert Hecht-Nielsen · 1989
Earlier work this paper cites.
Automatic differentiation and APL
Richard D. Neidinger · 1989
Earlier work this paper cites.
Untagged data in tagged environments: Choosing optimal representations at compile time
John Peterson · 1989
Earlier work this paper cites.
PADRE2, version 1—user’s manual
K. Kubo and M. Iri · 1990
Earlier work this paper cites.
MXYZPTLK: A practical, user-friendly C++ implementation of differential algebra: User’s guide
L. Michelotti · 1990
Earlier work this paper cites.
Extrapolation Methods: Theory and Practice
Claude Brezinski and M. Redivo Zaglia · 1991
Earlier work this paper cites.
Use of automatic differentiation for calculating Hessians and Newton steps
L. C. Dixon · 1991
Earlier work this paper cites.
Unboxed values as first class citizens in a non-strict functional language
Simon L Peyton Jones and John Launchbury · 1991
Earlier work this paper cites.
A taxonomy of automatic differentiation tools
David W. Juedes · 1991
Earlier work this paper cites.
Integration of automatic differentiation into a numerical library for PC’s
Vladimir Mazourik · 1991
Earlier work this paper cites.
Control-flow analysis of higher-order languages
Olin Shivers · 1991
Earlier work this paper cites.
Automatic differentiation in MATLAB
Lawrence C. Rich and David R. Hill · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Probabilistic inference using Markov chain Monte Carlo methods
Radford M. Neal · 1993
Earlier work this paper cites.
Reverse accumulation and attractive fixed points
Bruce Christianson · 1994
Earlier work this paper cites.
Parallel computation of automatic differentiation applied to magnetic field calculations
Ruth L. Hinkins · 1994
Earlier work this paper cites.
Fast exact multiplication by the Hessian
Barak A. Pearlmutter · 1994
Earlier work this paper cites.
Understanding the Metropolis-Hastings algorithm
Siddhartha Chib and Edward Greenberg · 1995
Earlier work this paper cites.
FADBAD, a flexible C++ package for automatic differentiation
Claus Bendtsen and Ole Stauning · 1996
Earlier work this paper cites.
Differential quadrature method in computational mechanics: A review
Charles W. Bert and Moinuddin Malik · 1996
Earlier work this paper cites.
COSY INFINITY and its applications in nonlinear dynamics
Martin Berz, Kyoko Makino, Khodr Shamseddine, Georg H. Hoffstätter, and Weishi Wan · 1996
Earlier work this paper cites.
ADIFOR 2.0: Automatic differentiation of Fortran 77 programs
Christian Bischof, Alan Carle, George Corliss, Andreas Griewank, and Paul Hovland · 1996
Earlier work this paper cites.
Numerical Methods for Unconstrained Optimization and Nonlinear Equations
John E. Dennis and Robert B. Schnabel · 1996
Earlier work this paper cites.
Automatically finding and exploiting partially separable structure in nonlinear programming problems
David M. Gay · 1996
Earlier work this paper cites.
ADIC: An extensible automatic differentiation tool for ANSI-C
Christian Bischof, Lucas Roh, and Andrew Mauer-Oats · 1997
Earlier work this paper cites.
Sensitivity analysis for atmospheric chemistry models via automatic differentiation
Gregory R. Carmichael and Adrian Sandu · 1997
Earlier work this paper cites.
Generative models for discovering sparse distributed representations
Geoffrey E. Hinton and Zoubin Ghahramani · 1997
Earlier work this paper cites.
Automatic differentiation and interval arithmetic for estimation of disequilibrium models
Max E. Jerrell · 1997
Earlier work this paper cites.
The effectiveness of type-based unboxing
Xavier Leroy · 1997
Earlier work this paper cites.
Finite difference of adjoint or adjoint of finite difference?
Z. Sirkes and E. Tziperman · 1997
Earlier work this paper cites.
Algorithm 778: L-BFGS-B: Fortran subroutines for large-scale bound-constrained optimization
Ciyou Zhu, Richard H. Byrd, Peihuang Lu, and Jorge Nocedal · 1997
Earlier work this paper cites.
Online learning and stochastic approximations
Léon Bottou · 1998
Earlier work this paper cites.
Regularization tools for training large feed-forward neural networks using automatic differentiation
Jerry Eriksson, Mårten Gulliksson, Per Lindström, and Per Åke Wedin · 1998
Earlier work this paper cites.
Recipes for adjoint code construction
Ralf Giering and Thomas Kaminski · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Transformation invariance in pattern recognition, tangent distance and tangent propagation
Patrice Simard, Yann LeCun, John Denker, and Bernard Victorri · 1998
Earlier work this paper cites.
INTLAB—INTerval LABoratory
Siegfried M. Rump · 1999
Earlier work this paper cites.
Local gain adaptation in stochastic gradient descent
Nicol N. Schraudolph · 1999
Earlier work this paper cites.
Bundle adjustment—a modern synthesis
Bill Triggs, Philip F. McLauchlan, Richard I. Hartley, and Andrew W. Fitzgibbon · 1999
Earlier work this paper cites.
Efficient adjoint derivatives: Application to the meteorological model Meso-NH
Isabelle Charpentier and Mohammed Ghemires · 2000
Earlier work this paper cites.
An introduction to automatic differentiation
Arun Verma · 2000
Cited alongside, same era.
Numerical Analysis
Rirchard L. Burden and J. Douglas Faires · 2001
Cited alongside, same era.
Structure and Interpretation of Classical Mechanics
Gerald J. Sussman and Jack Wisdom · 2001
Cited alongside, same era.
Automatic differentiation for computational finance
Christian H. Bischof, H. Martin Bücker, and Bruno Lang · 2002
Cited alongside, same era.
Lush reference manual, 2002
Léon Bottou and Yann LeCun · 2002
Cited alongside, same era.
Application of automatic diffentiation to race car performance optimisation
Daniele Casanova, Robin S. Sharp, Mark Final, Bruce Christianson, and Pat Symonds · 2002
Cited alongside, same era.
Learning stochastic inverses
Andreas Stuhlmüller, Jacob Taylor, and Noah Goodman · 2013
Later among the works it cites.
The ADiMat handbook, 2013
J. Willkomm and A. Vehreschild · 2013
Later among the works it cites.
Grounded language learning from video described with sentences
Haonan Yu and Jeffrey Mark Siskind · 2013
Later among the works it cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Later among the works it cites.
A fast and accurate dependency parser using neural networks
Danqi Chen and Christopher Manning · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shaun A. Forth and Trevor P. Evans · 2002
Cited alongside, same era.
AMPL: A Modeling Language for Mathematical Programming
Robert Fourer, David M. Gay, and Brian W. Kernighan · 2002
Cited alongside, same era.
Optimal sizing of industrial structural mechanics problems using AD
Gundolf Haase, Ulrich Langer, Ewald Lindner, and Wolfram Mühlhuber · 2002
Cited alongside, same era.
Computer Algebra Handbook: Foundations, Applications, Systems
Johannes Grabmeier and Erich Kaltofen · 2003
Cited alongside, same era.
A mathematical view of automatic differentiation
Andreas Griewank · 2003
Cited alongside, same era.
Stochastic volatility: Bayesian computation using automatic differentiation and the extended Kalman filter
Renate Meyer, David A. Fournier, and Andreas Berg · 2003
Cited alongside, same era.
Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer · 2014
Later among the works it cites.
Amortized inference in probabilistic reasoning
Samuel Gershman and Noah Goodman · 2014
Later among the works it cites.
Probabilistic programming
Andrew D Gordon, Thomas A Henzinger, Aditya V Nori, and Sriram K Rajamani · 2014
Later among the works it cites.
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Later among the works it cites.
The no-U-turn sampler: Adaptively setting path lengths in Hamiltonian Monte Carlo
Matthew D. Hoffman and Andrew Gelman · 2014
Later among the works it cites.
Caffe: Convolutional architecture for fast feature embedding
Yangqing Jia, Evan Shelhamer, Jeff Donahue, Sergey Karayev, Jonathan Long, Ross Girshick, Sergio Guadarrama, and Trevor Darrell · 2014
Later among the works it cites.
Auto-encoding variational Bayes
Diederik P. Kingma and Max Welling · 2014
Later among the works it cites.
OpenDR: An approximate differentiable renderer
Matthew M. Loper and Michael J. Black · 2014
Later among the works it cites.
Stochastic backpropagation and approximate inference in deep generative models
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Later among the works it cites.
User-specific hand modeling from monocular depth sequences
Jonathan Taylor, Richard Stebbing, Varun Ramakrishna, Cem Keskin, Jamie Shotton, Shahram Izadi, Aaron Hertzmann, and Andrew Fitzgibbon · 2014
Later among the works it cites.
The Stan math library: Reverse-mode automatic differentiation in C++
Bob Carpenter, Matthew D Hoffman, Marcus Brubaker, Daniel Lee, Peter Li, and Michael Betancourt · 2015
Closest in time.
Binaryconnect: Training deep neural networks with binary weights during propagations
Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David · 2015
Closest in time.
Learning to transduce with unbounded memory
Edward Grefenstette, Karl Moritz Hermann, Mustafa Suleyman, and Phil Blunsom · 2015
Closest in time.
Deep learning with limited numerical precision
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan · 2015
Closest in time.
Caffe con troll: Shallow ideas to speed up deep learning
Stefan Hadjis, Firas Abuzaid, Ce Zhang, and Christopher Ré · 2015
Closest in time.
Inferring algorithmic patterns with stack-augmented recurrent nets
Armand Joulin and Tomas Mikolov · 2015
Closest in time.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Closest in time.
Picture: A probabilistic programming language for scene perception
Tejas D. Kulkarni, Pushmeet Kohli, Joshua B. Tenenbaum, and Vikash Mansinghka · 2015
Closest in time.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Closest in time.
Gradient-based hyperparameter optimization through reversible learning
Dougal Maclaurin, David Duvenaud, and Ryan Adams · 2015
Closest in time.
Markov chain Monte Carlo and variational inference: Bridging the gap
Tim Salimans, Diederik Kingma, and Max Welling · 2015
Closest in time.
Deep learning in neural networks: An overview
Jürgen Schmidhuber · 2015
Closest in time.
Akshay Srinivasan and Emanuel Todorov · 2015
Closest in time.
End-to-end memory networks
Sainbayar Sukhbaatar, Jason Weston, Rob Fergus, et al · 2015
Closest in time.
Chainer: a next-generation open source framework for deep learning
Seiya Tokui, Kenta Oono, Shohei Hido, and Justin Clayton · 2015
Closest in time.
Efficient and robust analysis-by-synthesis in vision: A computational framework, behavioral tests, and modeling neuronal representations
Ilker Yildirim, Tejas D. Kulkarni, Winrich A. Freiwald, and Joshua B. Tenenbaum · 2015
Closest in time.
TensorFlow: Large-scale machine learning on heterogeneous distributed systems
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, et al · 2016
Closest in time.
Second order stochastic optimization in linear time
Naman Agarwal, Brian Bullins, and Elad Hazan · 2016
Closest in time.
Diffsharp: An AD library for .NET languages
Atılım Güneş Baydin, Barak A. Pearlmutter, and Jeffrey Mark Siskind · 2016
Closest in time.
Atılım Güneş Baydin, Barak A. Pearlmutter, and Jeffrey Mark Siskind · 2016
Closest in time.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E. Curtis, and Jorge Nocedal · 2016
Closest in time.
Stan: A probabilistic programming language
Bob Carpenter, Andrew Gelman, Matt Hoffman, Daniel Lee, Ben Goodrich, Michael Betancourt, Michael A Brubaker, Jiqiang Guo, Peter Li, and Allen Riddell · 2016
Closest in time.
Attend, infer, repeat: Fast scene understanding with generative models
S. M. Ali Eslami, Nicolas Heess, Theophane Weber, Yuval Tassa, David Szepesvari, Koray Kavukcuoglu, and Geoffrey E. Hinton · 2016
Closest in time.
A primer on neural network models for natural language processing
Yoav Goldberg · 2016
Closest in time.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Closest in time.
Hybrid computing using a neural network with dynamic external memory
Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka Grabska-Barwińska, Sergio Gómez Colmenarejo, Edward Grefenstette, Tiago Ramalho, John Agapiou, et al · 2016
Closest in time.
Memory-efficient backpropagation through time
Audrunas Gruslys, Rémi Munos, Ivo Danihelka, Marc Lanctot, and Alex Graves · 2016
Closest in time.
Composing graphical models with neural networks for structured representations and fast inference
Matthew Johnson, David K Duvenaud, Alex Wiltschko, Ryan P Adams, and Sandeep R Datta · 2016
Closest in time.
Ask me anything: Dynamic memory networks for natural language processing
Ankit Kumar, Ozan Irsoy, Peter Ondruska, Mohit Iyyer, James Bradbury, Ishaan Gulrajani, Victor Zhong, Romain Paulus, and Richard Socher · 2016
Closest in time.
Modeling, Inference and Optimization with Composable Differentiable Procedures
Dougal Maclaurin · 2016
Closest in time.
Deep amortized inference for probabilistic programs
Daniel Ritchie, Paul Horsfall, and Noah D Goodman · 2016
Closest in time.
Probabilistic programming in Python using PyMC3
John Salvatier, Thomas V Wiecki, and Christopher Fonnesbeck · 2016
Closest in time.
CNTK: Microsoft’s open-source deep-learning toolkit
Frank Seide and Amit Agarwal · 2016
Closest in time.
Efficient implementation of a higher-order language with built-in AD
Jeffrey Mark Siskind and Barak A. Pearlmutter · 2016
Closest in time.
ADiJaC—Automatic differentiation of Java classfiles
Emil I. Slusanschi and Vlad Dumitrel · 2016
Closest in time.
A benchmark of selected algorithmic differentiation tools on some problems in machine learning and computer vision
Filip Srajer, Zuzana Kukelova, and Andrew Fitzgibbon · 2016
Closest in time.
Edward: A library for probabilistic modeling, inference, and criticism
Dustin Tran, Alp Kucukelbir, Adji B. Dieng, Maja Rudolph, Dawen Liang, and David M. Blei · 2016
Closest in time.
Learning simple algorithms from examples
Wojciech Zaremba, Tomas Mikolov, Armand Joulin, and Rob Fergus · 2016
Closest in time.
OptNet: Differentiable optimization as a layer in neural networks
Brandon Amos and J Zico Kolter · 2017
Closest in time.
Backpropagation through the void: Optimizing control variates for black-box gradient estimation
Will Grathwohl, Dami Choi, Yuhuai Wu, Geoff Roeder, and David Duvenaud · 2017
Closest in time.
Automatic differentiation variational inference
Alp Kucukelbir, Dustin Tran, Rajesh Ranganath, Andrew Gelman, and David M. Blei · 2017
Closest in time.
Inference compilation and universal probabilistic programming
Tuan Anh Le, Atılım Güneş Baydin, and Frank Wood · 2017
Closest in time.
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Closest in time.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Closest in time.
Learning disentangled representations with semi-supervised deep generative models
N. Siddharth, Brooks Paige, Jan-Willem van de Meent, Alban Desmaison, Noah D. Goodman, Pushmeet Kohli, Frank Wood, and Philip Torr · 2017
Closest in time.
Divide-and-conquer checkpointing for arbitrary programs with no user annotation
Jeffrey Mark Siskind and Barak A. Pearlmutter · 2017
Closest in time.
Felipe Petroski Such, Vashisht Madhavan, Edoardo Conti, Joel Lehman, Kenneth O. Stanley, and Jeff Clune · 2017
Closest in time.
Deep probabilistic programming
Dustin Tran, Matthew D. Hoffman, Rif A. Saurous, Eugene Brevdo, Kevin Murphy, and David M. Blei · 2017
Closest in time.
REBAR: Low-variance, unbiased gradient estimates for discrete latent variable models
George Tucker, Andriy Mnih, Chris J. Maddison, John Lawson, and Jascha Sohl-Dickstein · 2017
Closest in time.
Tangent: Automatic differentiation using source code transformation in Python
Bart van Merriënboer, Alexander B. Wiltschko, and Dan Moldovan · 2017
Closest in time.
Online learning rate adaptation with hypergradient descent
Atılım Güneş Baydin, Robert Cornish, David Martínez Rubio, Mark Schmidt, and Frank Wood · 2018
Closest in time.