Fetching the paper…
Reading the bibliography…
In this paper, we set forth a new vision of reinforcement learning developed by us over the past few years, one that yields mathematically rigorous solutions to longstanding important questions that have remained unresolved: (i) how to design reliable, convergent, and robust reinforcement learning algorithms (ii) how to guarantee that reinforcement learning satisfies pre-specified "safety" guarantees, and remains in a stable region of the parameter space (iii) how to design "off-policy" temporal difference learning algorithms in a reliable and stable manner, and finally (iv) how to integrate the study of reinforcement learning into the rich theory of stochastic optimization.
On the numerical solution of heat conduction problems in two and three space variables
J. Douglas and H. Rachford · 1956
Earlier work this paper cites.
Functions convexes duales et points proximaux dans un espace hilbertien
J. Moreau · 1962
Earlier work this paper cites.
On some nonlinear elliptic differential functional equations
P. Hartman and G. Stampacchia · 1966
Earlier work this paper cites.
The relaxation method of finding the common point of convex sets and its application to the solution of problems in convex programming
L. Bregman · 1967
Earlier work this paper cites.
The extragradient method for finding saddle points and other problems
G. M. Korpelevich · 1976
Earlier work this paper cites.
The extragradient method for finding saddle points and other problems
G. Korpelevich · 1977
Earlier work this paper cites.
Traffic equilibria and variational inequalities
S. Dafermos · 1980
Earlier work this paper cites.
Problem Complexity and Method Efficiency in Optimization
A. Nemirovksi and D. Yudin · 1983
Earlier work this paper cites.
Modification of the extragradient method for solving variational inequalities of certain optimization problems
E. Khobotov · 1987
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
R. S. Sutton · 1988
Earlier work this paper cites.
Learning quickly when irrelevant attributes abound: A new linear-threshold algorithm
N. Littlestone · 1988
Earlier work this paper cites.
Linear Complementarity, Linear and Nonlinear Programming
K. Murty · 1988
Earlier work this paper cites.
Learning from Delayed Rewards
C. Watkins · 1989
Earlier work this paper cites.
Migration equilibrium and variational inequalities
A. Nagurney · 1989
Earlier work this paper cites.
Application of Khobotov’s algorithm to variational inequalities and network equilibrium problems
P. Marcotte · 1991
Earlier work this paper cites.
Numerical Recipes in C
W. Press, S. Tuekolsky, W. Vettering, and B. Flannery · 1992
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T Polyak and A. B Juditsky · 1992
Earlier work this paper cites.
Markov Decision Processes
M. L. Puterman · 1994
Earlier work this paper cites.
On the worst-case analysis of temporal-difference learning algorithms
Robert Schapire and Manfred K. Warmuth · 1994
Earlier work this paper cites.
Exponentiated gradient versus gradient descent for linear predictors
J. Kivinen and M. K. Warmuth · 1995
Earlier work this paper cites.
PID Controllers: Theory, Design, and Tuning
K. J. Åström and T. Hägglund · 1995
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
L. C. Baird · 1995
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
S. Bradtke and A. Barto · 1996
Earlier work this paper cites.
Neuro-Dynamic Programming
D. Bertsekas and J. Tsitsiklis · 1996
Earlier work this paper cites.
Projected Dynamical Systems and Variational Inequalities with Applications
A. Nagurney and D. Zhang · 1996
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
R. Tibshirani · 1996
Earlier work this paper cites.
Exponentiated gradient methods for reinforcement learning
D. Precup and R. S. Sutton · 1997
Earlier work this paper cites.
Biped dynamic walking using reinforcement learning
H. Bendrahim and J. A. Franklin · 1997
Earlier work this paper cites.
Parallel and Distributed Computation: Numerical Methods
D. Bertsekas and J. Tsitsiklis · 1997
Earlier work this paper cites.
A variant of Korpelevich’s method for variational inequalities with a new search strategy
A. Iusem and B. Svaiter · 1997
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Why natural gradient?
S. Amari and S. Douglas · 1998
Earlier work this paper cites.
Natural gradient works efficiently in learning
S. Amari · 1998
Earlier work this paper cites.
Network Economics: A Variational Inequality Approach
A. Nagurney · 1999
Earlier work this paper cites.
The Theory of Learning in Games
D. Fudenberg and D. Levine · 1999
Earlier work this paper cites.
A new projection method for variational inequality problems
M. Solodov and B. Svaiter · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
Nash convergence of gradient dynamics in general-sum games
S. Singh, M. Kearns, and Y. Mansour · 2000
Earlier work this paper cites.
The ordered subsets mirror descent optimization method with applications to tomography
A. Ben-Tal, T. Margalit, and A. Nemirovski · 2001
Earlier work this paper cites.
Online learning control by association and reinforcement
J. Si and Y. Wang · 2001
Earlier work this paper cites.
A natural policy gradient
S. Kakade · 2002
Cited alongside, same era.
Multiagent learning with a variable learning rate
M. Bowling and M. Veloso · 2002
Cited alongside, same era.
Finite-Dimensional Variational Inequalities and Complimentarity Problems
F. Facchinei and Pang J · 2003
Cited alongside, same era.
Least-squares policy evaluation algorithms with linear function approximation
A. Nedic and D. Bertsekas · 2003
Cited alongside, same era.
Least-squares policy iteration
M. Lagoudakis and R. Parr · 2003
Cited alongside, same era.
Fast calculation of stabilizing PID controllers
M. T. Söylemez, N. Munro, and H. Baki · 2003
Cited alongside, same era.
Lyapunov design for safe reinforcement learning
Economics
P. Samuelson and W. Nordhaus · 2009
Later among the works it cites.
Projected equations, variational inequalities, and temporal difference methods
D. Bertsekas · 2009
Later among the works it cites.
Distributed optimization and statistical learning via the alternating direction method of multipliers
S. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein · 2010
Later among the works it cites.
Whole-Body Strategies for Mobility and Manipulation
P. Deegan · 2010
Later among the works it cites.
Algorithms for reinforcement learning
C. Szepesvári · 2010
Later among the works it cites.
Dual averaging methods for regularized stochastic learning and online optimization
L. Xiao · 2010
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. J. Perkins and A. G. Barto · 2003
Cited alongside, same era.
Convex Optimization
S. Boyd and L. Vandenberghe · 2004
Cited alongside, same era.
Introductory lectures on convex optimization: A basic course , volume 87
Y. Nesterov · 2004
Cited alongside, same era.
Prox-method with rate of convergence O(1/t) for variational inequalities with Lipschitz continuous monotone operators and smooth convex-concave saddle point problems
A. Nemirovski · 2005
Cited alongside, same era.
Non-Euclidean restricted memory level method for large-scale convex optimization
A. Ben-Tal and A. Nemirovski · 2005
Cited alongside, same era.
Representation policy iteration
S. Mahadevan · 2005
Cited alongside, same era.
GQ ( λ \lambda ): A general gradient algorithm for temporal-difference prediction learning with eligibility traces
H.R. Maei and R.S. Sutton · 2010
Later among the works it cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2010
Later among the works it cites.
A general framework for a class of first order primal-dual algorithms for convex optimization in imaging science
E. Esser, X. Zhang, and T. F Chan · 2010
Later among the works it cites.
LSTD with Random Projections
M. Ghavamzadeh, A. Lazaric, O. A. Maillard, and R. Munos · 2010
Later among the works it cites.
Networks, Crowds, and Markets: Reasoning About a Highly Connected World
D. Easley and J. Kleinberg · 2010
Later among the works it cites.
Fast active-set-type algorithms for L1-regularized linear regression
J. Kim and H. Park · 2010
Later among the works it cites.
Proximal splitting methods in signal processing
P. Combetes and J.C. Pesquel · 2011
Later among the works it cites.
Proximal splitting methods in signal processing
P. L Combettes and J. C Pesquet · 2011
Later among the works it cites.
Stochastic methods for l1 regularized loss minimization
S. Shalev-Shwartz and A. Tewari · 2011
Later among the works it cites.
Follow-the-regularized-leader and mirror descent: Equivalence theorems and l1 regularization
H. B. McMahan · 2011
Later among the works it cites.
Convex analysis and monotone operator theory in Hilbert spaces
H. H Bauschke and P. L Combettes · 2011
Later among the works it cites.
A first-order primal-dual algorithm for convex problems with applications to imaging
A. Chambolle and T. Pock · 2011
Later among the works it cites.
Finite-Sample Analysis of Lasso-TD
M. Ghavamzadeh, A. Lazaric, R. Munos, and M. Hoffman · 2011
Later among the works it cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Later among the works it cites.
Gradient temporal-difference learning algorithms
H. R. Maei · 2011
Later among the works it cites.
Value function approximation in reinforcement learning using the fourier basis
G. Konidaris, S. Osentoski, and P. S. Thomas · 2011
Later among the works it cites.
The Fixed Points of Off-Policy TD
J. Zico Kolter · 2011
Later among the works it cites.
An optimal method for stochastic composite optimization
G. Lan · 2012
Later among the works it cites.
Manifold identification in dual averaging for regularized stochastic online learning
S. Lee and S.J. Wright · 2012
Later among the works it cites.
Control design for Markov chains under safety constraints: A convex approach
E. Arvelo and N. C. Martins · 2012
Later among the works it cites.
Variational bayesian optimization for runtime risk-sensitive control
S. Kuindersma, R. Grupen, and A. G. Barto · 2012
Later among the works it cites.
Motor primitive discovery
P. S. Thomas and A. G. Barto · 2012
Later among the works it cites.
Model-free reinforcement learning with continuous action in practice
T. Degris, P. M. Pilarski, and R. S. Sutton · 2012
Later among the works it cites.
Bias in natural actor-critic algorithms
P. S. Thomas · 2012
Later among the works it cites.
A Dantzig Selector Approach to Temporal Difference Learning
M. Geist, B. Scherrer, A. Lazaric, and M. Ghavamzadeh · 2012
Later among the works it cites.
Static prediction games for adversarial learning problems
M. Bruckner, C. Kanzow, and T. Scheffer · 2012
Later among the works it cites.
A voted regularized dual averaging method for large-scale discriminative training in natural language processing
J. Gao, T. Xu, L. Xiao, and X. He · 2013
Later among the works it cites.
Proximal algorithms
N. Parikh and S. Boyd · 2013
Later among the works it cites.
Optimal primal-dual methods for a class of saddle point problems
Y. Chen, G. Lan, and Y. Ouyang · 2013
Later among the works it cites.
Practical kernel-based reinforcement learning
A. MS Barreto, D. Precup, and J. Pineau · 2013
Later among the works it cites.
Galerkin methods for complementarity problems and variational inequalities
G. Gordon · 2013
Later among the works it cites.
Sparse Reinforcement Learning via Convex Optimization
Z. Qin and W. Li · 2014
Closest in time.
Policy evaluation with temporal differences: A survey and comparison
C. Dann, G. Neumann, and J. Peters · 2014
Closest in time.