Fetching the paper…
Reading the bibliography…
Nonlinear control systems with partial information to the decision maker are prevalent in a variety of applications.
Applied Nonparametric Regression
Wolfgang Härdle · 1990
Earlier work this paper cites.
Kernel principal component analysis
Bernhard Schölkopf, Alexander Smola, and Klaus-Robert Müller · 1997
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Online convex optimization in the bandit setting: Gradient descent without a gradient
Abraham D. Flaxman, Adam Tauman Kalai, and H. Brendan McMahan · 2005
Earlier work this paper cites.
Optimal Control: Linear Quadratic Methods
Brian D.O. Anderson and John B. Moore · 2007
Earlier work this paper cites.
Support Vector Machines
Ingo Steinwart and Andreas Christmann · 2008
Earlier work this paper cites.
Recovering low-rank matrices from few coefficients in any basis
David Gross · 2011
Earlier work this paper cites.
Dynamic Programming and Optimal Control
Dimitri Bertsekas · 2012
Earlier work this paper cites.
Random feature maps for dot product kernels
Purushottam Kar and Harish Karnick · 2012
Earlier work this paper cites.
Introduction to the Non-asymptotic Analysis of Random Matrices
Roman Vershynin · 2012
Earlier work this paper cites.
Playing Atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Nonlinear Systems: Analysis, Stability, and Control
Shankar Sastry · 2013
Earlier work this paper cites.
Deep direct reinforcement learning for financial signal representation and trading
Yue Deng, Feng Bao, Youyong Kong, Zhiquan Ren, and Qionghai Dai · 2016
Earlier work this paper cites.
Regularized policy iteration with nonparametric function spaces
Amir-massoud Farahmand, Mohammad Ghavamzadeh, Csaba Szepesvári, and Shie Mannor · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Christopher J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Earlier work this paper cites.
Feedback linearization using Gaussian processes
Jonas Umlauft, Thomas Beckers, Melanie Kimmel, and Sandra Hirche · 2017
Cited alongside, same era.
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham Kakade, and Mehran Mesbahi · 2018
Cited alongside, same era.
LQR through the lens of first order methods: Discrete-time case
Jingjing Bu, Afshin Mesbahi, Maryam Fazel, and Mehran Mesbahi · 2019
Cited alongside, same era.
Reinforcement learning and deep learning based lateral control for autonomous driving [application notes]
Dong Li, Dongbin Zhao, Qichao Zhang, and Yaran Chen · 2019
Cited alongside, same era.
Neural trust region/proximal policy optimization attains globally optimal policy
Boyi Liu, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2019
Cited alongside, same era.
Neural policy gradient methods: Global optimality and rates of convergence
Lingxiao Wang, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2020
Later among the works it cites.
Feedback linearization for uncertain systems via reinforcement learning
Tyler Westenbroek, David Fridovich-Keil, Eric Mazumdar, Shreyas Arora, Valmik Prabhu, S. Shankar Sastry, and Claire J. Tomlin · 2020
Later among the works it cites.
Variational policy gradient method for reinforcement learning with general utilities
Junyu Zhang, Alec Koppel, Amrit Singh Bedi, Csaba Szepesvari, and Mengdi Wang · 2020
Later among the works it cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Alekh Agarwal, Sham M. Kakade, Jason Lee, and Gaurav Mahajan · 2021
Later among the works it cites.
On the linear convergence of policy gradient methods for finite MDPs
Jalaj Bhandari and Daniel Russo · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Derivative-free methods for policy optimization: Guarantees for linear quadratic systems
Dhruv Malik, Ashwin Pananjady, Kush Bhatia, Koulik Khamaru, Peter L. Bartlett, and Martin J. Wainwright · 2019
Cited alongside, same era.
Certainty equivalence is efficient for linear quadratic control
Horia Mania, Stephen Tu, and Benjamin Recht · 2019
Cited alongside, same era.
Global exponential convergence of gradient methods over the nonconvex landscape of the linear quadratic regulator
Hesameddin Mohammadi, Armin Zare, Mahdi Soltanolkotabi, and Mihailo R Jovanović · 2019
Cited alongside, same era.
Provably global convergence of actor-critic: A case for linear quadratic regulator with ergodic cost
Zhuoran Yang, Yongxin Chen, Mingyi Hong, and Zhaoran Wang · 2019
Cited alongside, same era.
Adaptive population extremal optimization-based PID neural network for multivariable nonlinear control systems
Guoqiang Zeng, Xiaoqing Xie, Minrong Chen, and Jian Weng · 2019
Cited alongside, same era.
On the sample complexity of the linear quadratic regulator
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu · 2020
Cited alongside, same era.
Natural policy gradient primal-dual method for constrained Markov decision processes
Dongsheng Ding, Kaiqing Zhang, Tamer Başar, and Mihailo R. Jovanović · 2020
Cited alongside, same era.
Fast global convergence of natural policy gradient methods with entropy regularization
Shicong Cen, Chen Cheng, Yuxin Chen, Yuting Wei, and Yuejie Chi · 2021
Later among the works it cites.
Single-timescale actor-critic provably finds globally optimal policy
Zuyue Fu, Zhuoran Yang, and Zhaoran Wang · 2021
Later among the works it cites.
Learning optimal controllers for linear systems with multiplicative noise via policy gradient
Benjamin Gravell, Peyman Mohajerin Esfahani, and Tyler Summers · 2021
Later among the works it cites.
Policy gradient methods for the noisy linear quadratic regulator over a finite horizon
Ben Hambly, Renyuan Xu, and Huining Yang · 2021
Later among the works it cites.
Exploiting linear models for model-free nonlinear control: A provably convergent policy gradient approach
Guannan Qu, Chenkai Yu, Steven Low, and Adam Wierman · 2021
Later among the works it cites.
Lukasz Szpruch, Tanut Treetanthiploet, and Yufei Zhang · 2021
Later among the works it cites.
Doubly robust off-policy actor-critic: Convergence and optimality
Tengyu Xu, Zhuoran Yang, Zhaoran Wang, and Yingbin Liang · 2021
Later among the works it cites.
Policy optimization for ℋ 2 \mathcal{H}_{2} linear control with ℋ ∞ \mathcal{H}_{\infty} robustness guarantee: Implicit regularization and global convergence
Kaiqing Zhang, Bin Hu, and Tamer Basar · 2021
Later among the works it cites.
Yufeng Zhang, Zhuoran Yang, and Zhaoran Wang · 2021
Later among the works it cites.
Lukasz Szpruch, Tanut Treetanthiploet, and Yufei Zhang · 2022
Later among the works it cites.
On the convergence rates of policy gradient methods
Lin Xiao · 2022
Later among the works it cites.
Stochastic policy gradient methods: Improved sample complexity for fisher-non-degenerate policies
Ilyas Fatkhullin, Anas Barakat, Anastasia Kireeva, and Niao He · 2023
Closest in time.