An Equivalence Between Adaptive Dynamic Programming With a Critic and Backpropagation Through Time
Michael Fairbank, Eduardo Alonso, and Daniel V Prokhorov · 2013
Later among the works it cites.
Gradient Weights help Nonparametric Regressors
Samory Kpotufe and Abdeslam Boularias · 2013
Later among the works it cites.
On Rectified Linear Units for Speech Processing
M D Zeiler, M Ranzato, R Monga, M Mao, K Yang, Q V Le, P Nguyen, A Senior, V Vanhoucke, J Dean, and G Hinton · 2013
Later among the works it cites.
Policy Evaluation with Temporal Differences: A Survey and Comparison
Christoph Dann, Gerhard Neumann, and Jan Peters · 2014
Later among the works it cites.
Doubly Robust Policy Evaluation and Optimization
Miroslav Dudík, Dumitru Erhan, John Langford, and Lihong Li · 2014
Later among the works it cites.
Combining Reward Shaping and Hierarchies for Scaling to Large Multiagent Systems
Chris HolmesParker, Adrian K Agogino, and Kagan Tumer · 2014
Later among the works it cites.
Deterministic Policy Gradient Algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Later among the works it cites.
Dropout: A Simple Way to Prevent Neural Networks from Overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Later among the works it cites.
A Consistent Estimator of the Expected Gradient Outerproduct
Shubhendu Trivedi, Jialei Wang, Samory Kpotufe, and Gregory Shakhnarovich · 2014
Later among the works it cites.
Deep Online Convex Optimization by Putting Forecaster to Sleep
Original
David Balduzzi · 2015
Closest in time.
Kickback cuts Backprop’s red-tape: Biologically plausible credit assignment in neural networks
David Balduzzi, Hastagiri Vanchinathan, and Joachim Buhmann · 2015
Closest in time.
End-to-End Training of Deep Visuomotor Policies
Original
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2015
Closest in time.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Closest in time.
From Pixels to Torques: Policy Learning with Deep Dynamical Models
Original
Niklas Wahlström, Thomas B. Schön, and Marc Peter Deisenroth · 2015
Closest in time.