Fetching the paper…
Reading the bibliography…
When function approximation is used, solving the Bellman optimality equation with stability guarantees has remained a major open problem in reinforcement learning for decades.
Generalized polynomial approximations in Markovian decision processes
Schweitzer, Paul J. and Seidmann, Abraham · 1985
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, Richard S · 1988
Earlier work this paper cites.
Learning from Delayed Rewards
Watkins, Christopher J.C.H · 1989
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, Long-Ji · 1992
Earlier work this paper cites.
On-line Q-learning using connectionist systems
Rummery, G. A. and Niranjan, M · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, Leemon · 1995
Earlier work this paper cites.
Generalization in reinforcement learning: Safely approximating the value function
Boyan, Justin A. and Moore, Andrew W · 1995
Earlier work this paper cites.
Stable function approximation in dynamic programming
Gordon, Geoffrey J · 1995
Earlier work this paper cites.
Sphere packing numbers for subsets of the Boolean n n -cube with bounded Vapnik-Chervonenkis dimension
Haussler, David · 1995
Earlier work this paper cites.
Neuro-Dynamic Programming
Bertsekas, Dimitri P. and Tsitsiklis, John N · 1996
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding
Sutton, Richard S · 1996
Earlier work this paper cites.
Stochastic approximation with two time scales
Vivek S Borkar · 1997
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
Tsitsiklis, John N. and Van Roy, Benjamin · 1997
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, Richard S and Barto, Andrew G · 1998
Earlier work this paper cites.
On the existence of fixed points for approximate value iteration and temporal-difference learning
de Farias, Daniela Pucci and Van Roy, Benjamin · 2000
Earlier work this paper cites.
Least-squares temporal difference learning
Boyan, Justin A · 2002
Earlier work this paper cites.
Mixing and moment properties of various GARCH and stochastic volatility models
Carrasco, Marine and Chen, Xiaohong · 2002
Earlier work this paper cites.
A natural policy gradient
Kakade, Sham · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, Michael J. and Singh, Satinder P · 2002
Earlier work this paper cites.
Kernel-based reinforcement learning
Ormoneit, Dirk and Sen, Śaunak · 2002
Earlier work this paper cites.
The linear programming approach to approximate dynamic programming
de Farias, Daniela Pucci and Van Roy, Benjamin · 2003
Earlier work this paper cites.
Least-squares policy iteration
Lagoudakis, Michail G. and Parr, Ronald · 2003
Earlier work this paper cites.
Convex Optimization
Boyd, Stephen and Vandenberghe, Lieven · 2004
Earlier work this paper cites.
Smooth minimization of non-smooth functions
Nesterov, Yu · 2005
Cited alongside, same era.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, Ronald J · 2005
Cited alongside, same era.
Linearly-solvable Markov decision problems
Todorov, Emanuel · 2006
Cited alongside, same era.
Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
Antos, András, Szepesvári, Csaba, and Munos, Rémi · 2008
Cited alongside, same era.
An analysis of reinforcement learning with function approximation
Melo, Francisco S., Meyn, Sean P., and Ribeiro, M. Isabel · 2008
Cited alongside, same era.
Robust stochastic approximation approach to stochastic programming
Nemirovski, Arkadi, Juditsky, Anatoli, Lan, Guanghui, and Shapiro, Alexander · 2009
Trust region policy optimization
Schulman, John, Levine, Sergey, Abbeel, Pieter, Jordan, Michael I, and Moritz, Philipp · 2015
Later among the works it cites.
Nonlinear Programming
Bertsekas, Dimitri P · 2016
Later among the works it cites.
Brockman, Greg, Cheung, Vicki, Pettersson, Ludwig, Schneider, Jonas, Schulman, John, Tang, Jie, and Zaremba, Wojciech · 2016
Later among the works it cites.
Stochastic primal-dual methods and sample complexity of reinforcement learning
Chen, Yichen and Wang, Mengdi · 2016
Later among the works it cites.
Learning from conditional distributions via dual embeddings
Bo Dai, Niao He, Yunpeng Pan, Byron Boots, and Le Song · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Lectures on Stochastic Programming: Modeling and Theory
Shapiro, Alexander, Dentcheva, Darinka, and Ruszczyński, Andrzej · 2009
Cited alongside, same era.
Reinforcement learning in finite MDPs: PAC analysis
Strehl, Alexander L., Li, Lihong, and Littman, Michael L · 2009
Cited alongside, same era.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Sutton, Richard S., Maei, Hamid Reza, Precup, Doina, Bhatnagar, Shalabh, Silver, David, Szepesvári, Csaba, and Wiewiora, Eric · 2009
Cited alongside, same era.
Toward off-policy learning control with function approximation
Maei, Hamid Reza, Szepesvári, Csaba, Bhatnagar, Shalabh, and Sutton, Richard S · 2010
Cited alongside, same era.
Gradient Temporal-Difference Learning Algorithms
Maei, Hamid Reza · 2011
Cited alongside, same era.
On stochastic optimal control and reinforcement learning by approximate inference
Rawlik, Konrad, Toussaint, Marc, and Vijayakumar, Sethu · 2012
Cited alongside, same era.
Taming the noise in reinforcement learning via soft updates
Fox, Roy, Pakman, Ari, and Tishby, Naftali · 2016
Later among the works it cites.
Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
Ghadimi, Saeed, Lan, Guanghui, and Zhang, Hongchao · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Mnih, Volodymyr, Badia, Adrià Puigdomènech, Mirza, Mehdi, Graves, Alex, Lillicrap, Timothy P., Harley, Tim, Silver, David, and Kavukcuoglu, Koray · 2016
Later among the works it cites.
An alternative softmax operator for reinforcement learning
Asadi, Kavosh and Littman, Michael L · 2017
Closest in time.
Stochastic variance reduction methods for policy evaluation
Du, Simon S., Chen, Jianshu, Li, Lihong, Xiao, Lin, and Zhou, Dengyong · 2017
Closest in time.
Reinforcement learning with deep energy-based policies
Haarnoja, Tuomas, Tang, Haoran, Abbeel, Pieter, and Levine, Sergey · 2017
Closest in time.
Gans trained by a two time-scale update rule converge to a nash equilibrium
Heusel, Martin, Ramsauer, Hubert, Unterthiner, Thomas, Nessler, Bernhard, Klambauer, Günter, and Hochreiter, Sepp · 2017
Closest in time.
Stein variational policy gradient
Liu, Yang, Ramachandran, Prajit, Liu, Qiang, and Peng, Jian · 2017
Closest in time.
Bridging the gap between value and policy based reinforcement learning
Nachum, Ofir, Norouzi, Mohammad, Xu, Kelvin, and Schuurmans, Dale · 2017
Closest in time.
A unified view of entropy-regularized markov decision processes, 2017
Neu, Gergely, Jonsson, Anders, and Gómez, Vicenç · 2017
Closest in time.
Towards generalization and simplicity in continuous control
Rajeswaran, Aravind, Lowrey, Kendall, Todorov, Emanuel V., and Kakade, Sham M · 2017
Closest in time.
Equivalence between policy gradients and soft Q-learning, 2017
Schulman, John, Abbeel, Pieter, and Chen, Xi · 2017
Closest in time.
Randomized linear programming solves the discounted Markov decision problem in nearly-linear running time
Wang, Mengdi · 2017
Closest in time.
Boosting the actor with dual critic
Dai, Bo, Shaw, Albert, He, Niao, Li, Lihong, and Song, Le · 2018
Closest in time.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, Tuomas, Zhou, Aurick, Abbeel, Pieter, and Levine, Sergey · 2018
Closest in time.
Deep reinforcement learning that matters
Henderson, Peter, Islam, Riashat, Bachman, Philip, Pineau, Joelle, Precup, Doina, and Meger, David · 2018
Closest in time.
Trust-PCL: An off-policy trust region method for continuous control
Nachum, Ofir, Norouzi, Mohammad, Xu, Kelvin, and Schuurmans, Dale · 2018
Closest in time.