Fetching the paper…
Reading the bibliography…
We propose a new policy gradient method, named homotopic policy mirror descent (HPMD), for solving discounted, infinite horizon MDPs with finite state and action spaces.
Problem complexity and method efficiency in optimization
Arkadij Semenovič Nemirovskij and David Borisovich Yudin · 1983
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Amir Beck and Marc Teboulle · 2003
Earlier work this paper cites.
The entire regularization path for the support vector machine
Trevor Hastie, Saharon Rosset, Robert Tibshirani, and Ji Zhu · 2004
Earlier work this paper cites.
Boosting as a regularized path to a maximum margin classifier
Saharon Rosset, Ji Zhu, and Trevor Hastie · 2004
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2005
Earlier work this paper cites.
Numerical optimization
Jorge Nocedal and Stephen Wright · 2006
Earlier work this paper cites.
L1-regularization path algorithm for generalized linear models
Mee Young Park and Trevor Hastie · 2007
Earlier work this paper cites.
Peng Zhao and Bin Yu · 2007
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
Relative entropy policy search
Jan Peters, Katharina Mulling, and Yasemin Altun · 2010
Earlier work this paper cites.
The simplex and policy-iteration methods are strongly polynomial for the markov decision problem with a fixed discount rate
Yinyu Ye · 2011
Cited alongside, same era.
Improved and generalized upper bounds on the complexity of policy iteration
Bruno Scherrer · 2013
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
A unified view of entropy-regularized markov decision processes
Gergely Neu, Anders Jonsson, and Vicenç Gómez · 2017
Cited alongside, same era.
Implicit bias of gradient descent based adversarial training on separable data
Yan Li, Ethan X.Fang, Huan Xu, and Tuo Zhao · 2020
Later among the works it cites.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps
Lior Shani, Yonathan Efroni, and Shie Mannor · 2020
Later among the works it cites.
Fast global convergence of natural policy gradient methods with entropy regularization
Shicong Cen, Chen Cheng, Yuxin Chen, Yuting Wei, and Yuejie Chi · 2021
Later among the works it cites.
Twice regularized mdps and the equivalence between robustness and regularization
Esther Derman, Matthieu Geist, and Shie Mannor · 2021
Later among the works it cites.
Actor-critic is implicitly biased towards high entropy optimal policies
Yuzheng Hu, Ziwei Ji, and Matus Telgarsky · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Cited alongside, same era.
The implicit bias of gradient descent on nonseparable data
Ziwei Ji and Matus Telgarsky · 2019
Cited alongside, same era.
Neural trust region/proximal policy optimization attains globally optimal policy
Boyi Liu, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2019
Cited alongside, same era.
Optimality and approximation with policy gradient methods in markov decision processes
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2020
Cited alongside, same era.
A note on the linear convergence of policy gradient methods
Jalaj Bhandari and Daniel Russo · 2020
Cited alongside, same era.
Sajad Khodadadian, Prakirt Raj Jhunjhunwala, Sushil Mahavir Varma, and Siva Theja Maguluri · 2021
Later among the works it cites.
Implicit regularization of bregman proximal point algorithm and mirror descent on separable data
Yan Li, Caleb Ju, Ethan X Fang, and Tuo Zhao · 2021
Later among the works it cites.
Wenhao Zhan, Shicong Cen, Baihe Huang, Yuxin Chen, Jason D Lee, and Yuejie Chi · 2021
Later among the works it cites.
Policy mirror descent for reinforcement learning: Linear convergence, new sampling complexity, and generalized problem classes
Guanghui Lan · 2022
Closest in time.
Guanghui Lan, Yan Li, and Tuo Zhao · 2022
Closest in time.
On the convergence rates of policy gradient methods
Lin Xiao · 2022
Closest in time.