Fetching the paper…
Reading the bibliography…
Deterministic Policy Gradient (DPG) removes a level of randomness from standard randomized-action Policy Gradient (PG), and demonstrates substantial empirical success for tackling complex dynamic problems involving Markov decision processes.
Temporal difference learning and td-gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Inventory management in supply chains: a reinforcement learning approach
Ilaria Giannoccaro and Pierpaolo Pontrandolfo · 2002
Earlier work this paper cites.
Onactor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2003
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Evan Greensmith, Peter L Bartlett, and Jonathan Baxter · 2004
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters · 2013
Earlier work this paper cites.
Mobi-cliques for improving ergodic secrecy in fading wiretap channels under power constraints
Dionysios S Kalogerias and Athina P Petropulu · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Earlier work this paper cites.
Optimal rates for zero-order convex optimization: The power of two function evaluations
John C Duchi, Michael I Jordan, Martin J Wainwright, and Andre Wibisono · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Policy gradient in lipschitz markov decision processes
Matteo Pirotta, Marcello Restelli, and Luca Bascetta · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
Saeed Ghadimi, Guanghui Lan, and Hongchao Zhang · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Deep reinforcement learning: A brief survey
Kai Arulkumaran, Marc Peter Deisenroth, Miles Brundage, and Anil Anthony Bharath · 2017
Cited alongside, same era.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
Stochastic policy gradient ascent in reproducing kernel hilbert spaces
Santiago Paternain, Juan Andrés Bazerque, Austin Small, and Alejandro Ribeiro · 2018
Later among the works it cites.
Learning optimal resource allocations in wireless systems
Mark Eisen, Clark Zhang, Luiz FO Chamon, Daniel D Lee, and Alejandro Ribeiro · 2019
Later among the works it cites.
Model-free learning of optimal ergodic policies in wireless systems
Dionysios S Kalogerias, Mark Eisen, George J Pappas, and Alejandro Ribeiro · 2019
Later among the works it cites.
Zeroth-order stochastic compositional algorithms for risk-aware learning
Dionysios S Kalogerias and Warren B Powell · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shixiang Gu, Ethan Holly, Timothy Lillicrap, and Sergey Levine · 2017
Cited alongside, same era.
Random gradient-free minimization of convex functions
Yurii Nesterov and Vladimir Spokoiny · 2017
Cited alongside, same era.
A finite time analysis of temporal difference learning with linear function approximation
Jalaj Bhandari, Daniel Russo, and Raghav Singal · 2018
Cited alongside, same era.
Finite sample analyses for td (0) with function approximation
Gal Dalal, Balázs Szörényi, Gugan Thoppe, and Shie Mannor · 2018
Cited alongside, same era.
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham M Kakade, and Mehran Mesbahi · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Van Hoof, and David Meger · 2018
Cited alongside, same era.
Stochastic variance-reduced policy gradient
Matteo Papini, Damiano Binaghi, Giuseppe Canonaco, Matteo Pirotta, and Marcello Restelli · 2018
Cited alongside, same era.
Harshat Kumar, Alec Koppel, and Alejandro Ribeiro · 2019
Later among the works it cites.
Hesameddin Mohammadi, Armin Zare, Mahdi Soltanolkotabi, and Mihailo R Jovanović · 2019
Later among the works it cites.
Contrasting exploration in parameter and action space: A zeroth-order optimization perspective
Anirudh Vemula, Wen Sun, and J Andrew Bagnell · 2019
Later among the works it cites.
Reinforcement leaning in feature space: Matrix bandit, kernels, and regret bound
Lin F Yang and Mengdi Wang · 2019
Later among the works it cites.
Deep reinforcement learning based resource allocation for v2v communications
Hao Ye, Geoffrey Ye Li, and Biing-Hwang Fred Juang · 2019
Later among the works it cites.
Global convergence of policy gradient methods to (almost) locally optimal policies
Kaiqing Zhang, Alec Koppel, Hao Zhu, and Tamer Başar · 2019
Later among the works it cites.
Derivative-free methods for policy optimization: Guarantees for linear quadratic systems
Dhruv Malik, Ashwin Pananjady, Kush Bhatia, Koulik Khamaru, Peter L Bartlett, and Martin J Wainwright · 2020
Closest in time.