Fetching the paper…
Reading the bibliography…
Policy optimization (PO) is a key ingredient for reinforcement learning (RL).
On topological and metrical properties of stabilizing feedback gains: the MIMO case
1904
Earlier work this paper cites.
An improved convergence analysis of stochastic variance-reduced policy gradient
1905
Earlier work this paper cites.
Global optimality guarantees for policy gradient methods
1906
Earlier work this paper cites.
1907
Earlier work this paper cites.
Logarithmic regret for online control
1909
Earlier work this paper cites.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized MDPs
1909
Earlier work this paper cites.
Global convergence of policy gradient for sequential zero-sum linear quadratic dynamic games
1911
Earlier work this paper cites.
Gradient methods for minimizing functionals
1963
Earlier work this paper cites.
On the input-output stability of time-varying nonlinear feedback systems part one: Conditions derived using concepts of loop gain, conicity, and positivity
1966
Earlier work this paper cites.
On an iterative technique for Riccati equation computations
1968
Earlier work this paper cites.
An iterative technique for the computation of the steady state gains for the discrete optimal regulator
1971
Earlier work this paper cites.
Optimal stochastic linear systems with exponential performance criteria and their relation to deterministic differential games
1973
Earlier work this paper cites.
The stabilizing solution of the algebraic Riccati equation
1973
Earlier work this paper cites.
On the Lyapunov matrix equation
1974
Earlier work this paper cites.
On the Goldstein-Levitin-Polyak gradient projection method
1976
Earlier work this paper cites.
Risk-sensitive linear/quadratic/Gaussian control
1981
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate
1983
Earlier work this paper cites.
Matrix differential calculus with applications to simple, Hadamard, and Kronecker products
1985
Earlier work this paper cites.
Computational methods for parametric LQ problems–A survey
1987
Earlier work this paper cites.
A new CAD method and associated architectures for linear controllers
1988
Earlier work this paper cites.
State-space formulae for all stabilizing controllers that satisfy an
1988
Earlier work this paper cites.
Existence and comparison theorems for algebraic Riccati equations for continuous- and discrete-time systems
1988
Earlier work this paper cites.
LQG control with an
1989
Earlier work this paper cites.
State-space solutions to standard
1989
Earlier work this paper cites.
Relations between maximum-entropy/
1989
Earlier work this paper cites.
Minimum entropy
1990
Earlier work this paper cites.
Risk-sensitive Optimal Control
1990
Earlier work this paper cites.
Mixed-norm
1991
Earlier work this paper cites.
LQG cost bounds in discrete-time
1991
Earlier work this paper cites.
Risk sensitive optimal control and differential games
1992
Earlier work this paper cites.
On computing the stabilizing solution of the discrete-time Riccati equation
1992
Earlier work this paper cites.
Linear Matrix Inequalities in System and Control Theory
1994
Earlier work this paper cites.
Adaptive linear quadratic control using policy iteration
1994
Earlier work this paper cites.
A linear matrix inequality approach to
1994
Earlier work this paper cites.
The discrete-time Riccati equation related to the
1994
Earlier work this paper cites.
H-infinity Optimal Control and Related Minimax Design Problems: A Dynamic Game Approach
1995
Earlier work this paper cites.
A linear matrix inequality approach to the general mixed
1995
Earlier work this paper cites.
Multiobjective h/sub 2//h/sub/spl infin//control
1995
Earlier work this paper cites.
Analysis of discrete-time linear periodic systems
1996
Earlier work this paper cites.
h ∞ h_{\infty} design with pole placement constraints: an LMI approach
1996
Earlier work this paper cites.
On the Kalman-Yakubovich-Popov Lemma
1996
Earlier work this paper cites.
Robust and Optimal Control
1996
Cited alongside, same era.
Risk-sensitive control of finite state machines on an infinite horizon i
1997
Cited alongside, same era.
Computational design of optimal output feedback controllers
1997
Cited alongside, same era.
Multiobjective
1998
Cited alongside, same era.
An exact solution to general four-block discrete-time mixed
1998
Cited alongside, same era.
Constrained Markov decision processes
1999
Cited alongside, same era.
Actor-critic algorithms
2000
Cited alongside, same era.
2015
Later among the works it cites.
Risk-sensitive and robust decision-making: A CVaR optimization approach
2015
Later among the works it cites.
Continuous control with deep reinforcement learning
2015
Later among the works it cites.
End-to-end training of deep visuomotor policies
2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Policy gradient methods for reinforcement learning with function approximation
2000
Cited alongside, same era.
Risk-sensitive optimal control for Markov decision processes with monotone cost
2002
Cited alongside, same era.
Convergence guarantees of policy optimization methods for markovian jump linear systems
2002
Cited alongside, same era.
A natural policy gradient
2002
Cited alongside, same era.
Switching in Systems and Control
2003
Cited alongside, same era.
2016
Later among the works it cites.
Constrained policy optimization
2017
Later among the works it cites.
Safe model-based reinforcement learning with stability guarantees
2017
Later among the works it cites.
2017
Later among the works it cites.
Risk-constrained reinforcement learning with percentile risk criteria
2017
Later among the works it cites.
On the sample complexity of the linear quadratic regulator
2017
Later among the works it cites.
Implicit regularization in nonconvex statistical estimation: Gradient descent converges linearly for phase retrieval, matrix completion and blind deconvolution
2017
Later among the works it cites.
Random gradient-free minimization of convex functions
2017
Later among the works it cites.
Robust adversarial reinforcement learning
2017
Later among the works it cites.
Proximal policy optimization algorithms
2017
Later among the works it cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
2018
Later among the works it cites.
A Lyapunov-based approach to safe reinforcement learning
2018
Later among the works it cites.
Online linear quadratic control
2018
Later among the works it cites.
Global convergence of policy gradient methods for the linear quadratic regulator
2018
Later among the works it cites.
Convergence guarantees for a class of non-convex and non-smooth optimization problems
2018
Later among the works it cites.
Openai five
2018
Later among the works it cites.
Stochastic variance-reduced policy gradient
2018
Later among the works it cites.
2018
Later among the works it cites.
Convergence analysis of gradient-based learning with non-uniform learning rates in non-cooperative multi-agent settings
2019
Closest in time.
Implicit regularization in over-parameterized neural networks
2019
Closest in time.
Rapid, robust, and reliable blind deconvolution via nonconvex optimization
2019
Closest in time.
Kernel-based reinforcement learning in robust markov decision processes
2019
Closest in time.
Derivative-free methods for policy optimization: Guarantees for linear quadratic systems
2019
Closest in time.
Robust reinforcement learning for continuous control with model misspecification
2019
Closest in time.
From self-tuning regulators to reinforcement learning and back again
2019
Closest in time.
Policy-gradient algorithms have no guarantees of convergence in continuous action and state multi-agent settings
2019
Closest in time.
A tour of reinforcement learning: The view from continuous control
2019
Closest in time.
On the convergence to stationary points of the iterative linear exponential quadratic Gaussian algorithm
2019
Closest in time.
Action robust reinforcement learning and applications in continuous control
2019
Closest in time.
Recovering robustness in model-free reinforcement learning
2019
Closest in time.
AlphaStar: Mastering the Real-Time Strategy Game StarCraft II
2019
Closest in time.
Neural policy gradient methods: Global optimality and rates of convergence
2019
Closest in time.
On the global convergence of actor-critic: A case for linear quadratic regulator with ergodic cost
2019
Closest in time.
Convergent policy optimization for safe reinforcement learning
2019
Closest in time.
Derivative-free policy optimization for risk-sensitive and robust control design: Implicit regularization and sample complexity
2021
Closest in time.