Fetching the paper…
Reading the bibliography…
Deterministic-policy actor-critic algorithms for continuous control improve the actor by plugging its actions into the critic and ascending the action-value gradient, which is obtained by chaining the actor's Jacobian matrix with the gradient of the critic with respect to input actions.
Advanced forecasting methods for global crisis warning and models of intelligence
Paul Werbos · 1977
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Making the world differentiable: On using self-supervised fully recurrent neural networks for dynamic reinforcement learning and planning in non-stationary environments
Jürgen Schmidhuber · 1990
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Richard S Sutton · 1991
Earlier work this paper cites.
Improving generalization performance using double backpropagation
Harris Drucker and Yann Le Cun · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Penalty functions
Alice E Smith, David W Coit, Thomas Baeck, David Fogel, and Zbigniew Michalewicz · 1995
Earlier work this paper cites.
Adaptive critic designs
Danil V Prokhorov and Donald C Wunsch · 1997
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
John N Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Jonathan Baxter and Peter L Bartlett · 2001
Earlier work this paper cites.
Model-based policy gradient reinforcement learning
Xin Wang and Thomas G Dietterich · 2003
Earlier work this paper cites.
Using inaccurate models in reinforcement learning
Pieter Abbeel, Morgan Quigley, and Andrew Y Ng · 2006
Earlier work this paper cites.
Reinforcement learning by value gradients
Michael Fairbank · 2008
Earlier work this paper cites.
Contractive auto-encoders: Explicit invariance during feature extraction
Salah Rifai, Pascal Vincent, Xavier Muller, Xavier Glorot, and Yoshua Bengio · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
A survey on policy search for robotics
Marc Peter Deisenroth, Gerhard Neumann, Jan Peters, et al · 2013
Earlier work this paper cites.
Reinforcement learning with misspecified model classes
Joshua Joseph, Alborz Geramifard, John W Roberts, Jonathan P How, and Nicholas Roy · 2013
Earlier work this paper cites.
Value-gradient learning
Michael Fairbank · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
Model-based policy gradients with parameter-based exploration by least-squares conditional density estimation
Voot Tangkaratt, Syogo Mori, Tingting Zhao, Jun Morimoto, and Masashi Sugiyama · 2014
Cited alongside, same era.
Compatible value gradients for reinforcement learning of continuous deep policies
David Balduzzi and Muhammad Ghifary · 2015
Cited alongside, same era.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Timothy Lillicrap, Tom Erez, and Yuval Tassa · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Model-based value estimation for efficient model-free reinforcement learning
Vladimir Feinberg, Alvin Wan, Ion Stoica, Michael I Jordan, Joseph E Gonzalez, and Sergey Levine · 2018
Later among the works it cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Later among the works it cites.
Model-ensemble trust-region policy optimization
Thanard Kurutach, Ignasi Clavera, Yan Duan, Aviv Tamar, and Pieter Abbeel · 2018
Later among the works it cites.
Adversarial regularizers in inverse problems
Sebastian Lunz, Ozan Öktem, and Carola-Bibiane Schönlieb · 2018
Later among the works it cites.
Policy optimization via importance sampling
Alberto Maria Metelli, Matteo Papini, Francesco Faccio, and Marcello Restelli · 2018
Later among the works it cites.
Reinforcement learning: An introduction
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep learning in neural networks: An overview
Jürgen Schmidhuber · 2015
Cited alongside, same era.
Gradient estimation using stochastic computation graphs
John Schulman, Nicolas Heess, Theophane Weber, and Pieter Abbeel · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
The concrete distribution: A continuous relaxation of discrete random variables
Chris J Maddison, Andriy Mnih, and Yee Whye Teh · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Sobolev training for neural networks
Wojciech M Czarnecki, Simon Osindero, Max Jaderberg, Grzegorz Swirszcz, and Razvan Pascanu · 2017
Cited alongside, same era.
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Efficient and accurate estimation of lipschitz constants for deep neural networks
Mahyar Fazlyab, Alexander Robey, Hamed Hassani, Manfred Morari, and George Pappas · 2019
Later among the works it cites.
DeepMDP: Learning continuous latent space models for representation learning
Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, and Marc G. Bellemare · 2019
Later among the works it cites.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Later among the works it cites.
On the variance of the adaptive learning rate and beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
On approximating ∇ f \nabla f with neural networks
Saeed Saremi · 2019
Later among the works it cites.
Neural empirical bayes
Saeed Saremi and Aapo Hyvarinen · 2019
Later among the works it cites.
Model-based active exploration
Pranav Shyam, Wojciech Jaśkowski, and Faustino Gomez · 2019
Later among the works it cites.
First-order adversarial vulnerability of neural networks and input dimension
Carl-Johann Simon-Gabriel, Yann Ollivier, Leon Bottou, Bernhard Schölkopf, and David Lopez-Paz · 2019
Later among the works it cites.
When to use parametric models in reinforcement learning?
Hado P van Hasselt, Matteo Hessel, and John Aslanides · 2019
Later among the works it cites.
Credit assignment techniques in stochastic computation graphs
Théophane Weber, Nicolas Heess, Lars Buesing, and David Silver · 2019
Later among the works it cites.
Model-augmented actor-critic: Backpropagating through paths
Ignasi Clavera, Yao Fu, and Pieter Abbeel · 2020
Closest in time.
Gradient-aware model-based policy search
Pierluca D’Oro, Alberto Maria Metelli, Andrea Tirinzoni, Matteo Papini, and Marcello Restelli · 2020
Closest in time.
A closer look at deep policy gradients
Andrew Ilyas, Logan Engstrom, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, and Aleksander Madry · 2020
Closest in time.