Fetching the paper…
Reading the bibliography…
This paper develops the first policy gradient method with global optimality guarantee and complexity analysis for robust reinforcement learning under model mismatch.
A topological property of real analytic subsets
Lojasiewicz, S · 1963
Earlier work this paper cites.
Gradient methods for minimizing functionals
Polyak, B. T · 1963
Earlier work this paper cites.
A robust version of the probability ratio test
Huber, P. J · 1965
Earlier work this paper cites.
Markovian decision processes with uncertain transition probabilities
Satia, J. K. and Lave Jr, R. E · 1973
Earlier work this paper cites.
Optimization and nonsmooth analysis
Clarke, F. H · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
On the generation of Markov decision processes
Archibald, T., McKinnon, K., and Thomas, L · 1995
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., Mansour, Y., et al · 1999
Earlier work this paper cites.
The ode method for convergence of stochastic approximation and reinforcement learning
Borkar, V. S. and Meyn, S. P · 2000
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N · 2000
Earlier work this paper cites.
Solving uncertain Markov decision processes
Bagnell, J. A., Ng, A. Y., and Schneider, J. G · 2001
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M · 2001
Earlier work this paper cites.
Complexity of finding stationary points of nonsmooth nonconvex functions
Zhang, J., Lin, H., Jegelka, S., Jadbabaie, A., and Sra, S · 2002
Earlier work this paper cites.
Nonparametric representation of policies and value functions: A trajectory-based approach
Atkeson, C. G. and Morimoto, J · 2003
Earlier work this paper cites.
On Fréchet subdifferentials
Kruger, A. Y · 2003
Earlier work this paper cites.
Robustness in Markov decision problems with uncertain transition matrices
Nilim, A. and El Ghaoui, L · 2004
Earlier work this paper cites.
Search and knightian uncertainty
Nishimura, K. G. and Ozaki, H · 2004
Earlier work this paper cites.
Robust dynamic programming
Iyengar, G. N · 2005
Earlier work this paper cites.
Robust reinforcement learning
Morimoto, J. and Doya, K · 2005
Earlier work this paper cites.
An axiomatic approach to ϵ \epsilon -contamination
Nishimura, K. G. and Ozaki, H · 2006
Earlier work this paper cites.
Stochastic approximation: a dynamical systems viewpoint , volume 48
Borkar, V. S · 2009
Earlier work this paper cites.
Robust Statistics
Huber, P. and Ronchetti, E · 2009
Earlier work this paper cites.
Distributionally robust markov decision processes
Xu, H. and Mannor, S · 2010
Earlier work this paper cites.
Robust modified policy iteration
Kaufman, D. L. and Schaefer, A. J · 2013
Earlier work this paper cites.
Reinforcement learning in robust Markov decision processes
Lim, S. H., Xu, H., and Mannor, S · 2013
Earlier work this paper cites.
Robust Markov decision processes
Wiesemann, W., Kuhn, D., and Rustem, B · 2013
Earlier work this paper cites.
Geometric measure theory
Federer, H · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
Scaling up robust mdps using function approximation
Tamar, A., Mannor, S., and Xu, H · 2014
Earlier work this paper cites.
On TD(0) with function approximation: Concentration bounds and a centered variant with exponential convergence
Korda, N. and La, P · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
Distributionally robust counterpart in markov decision processes
Yu, P. and Xu, H · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
Ghadimi, S. and Lan, G · 2016
Cited alongside, same era.
Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
Wasserstein robust reinforcement learning
Abdullah, M. A., Ren, H., Ammar, H. B., Milenkovic, V., Luo, R., Zhang, M., and Wang, J · 2019
Later among the works it cites.
Global optimality guarantees for policy gradient methods
Bhandari, J. and Russo, D · 2019
Later among the works it cites.
Neural temporal-difference learning converges to global optima
Cai, Q., Yang, Z., Lee, J. D., and Wang, Z · 2019
Later among the works it cites.
Gradient descent finds global minima of deep neural networks
Du, S., Lee, J., Li, H., Wang, L., and Zhai, X · 2019
Later among the works it cites.
Kernel-based reinforcement learning in robust markov decision processes
Lim, S. H. and Autef, A · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ghadimi, S., Lan, G., and Zhang, H · 2016
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the Polyak-łojasiewicz condition
Karimi, H., Nutini, J., and Schmidt, M · 2016
Cited alongside, same era.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P · 2017
Cited alongside, same era.
An alternative softmax operator for reinforcement learning
Asadi, K. and Littman, M. L · 2017
Cited alongside, same era.
First-order methods in optimization
Beck, A · 2017
Cited alongside, same era.
Adversarial attacks on neural network policies
Huang, S., Papernot, N., Goodfellow, I., Duan, Y., and Abbeel, P · 2017
Cited alongside, same era.
Delving into adversarial attacks on deep policies
Kos, J. and Song, D · 2017
Cited alongside, same era.
Action robust reinforcement learning and applications in continuous control
Tessler, C., Efroni, Y., and Mannor, S · 2019
Later among the works it cites.
Robust reinforcement learning with wasserstein constraint
Hou, L., Pang, L., Hong, X., Lan, Y., Ma, Z., and Yin, D · 2020
Later among the works it cites.
On the global convergence rates of softmax policy gradient methods
Mei, J., Xiao, C., Szepesvari, C., and Schuurmans, D · 2020
Later among the works it cites.
Robust constrained-MDPs: Soft-constrained robust policy optimization under model uncertainty
Russel, R. H., Benosman, M., and Van Baar, J · 2020
Later among the works it cites.
Convergence of a stochastic subgradient method with averaging for nonsmooth nonconvex constrained optimization
Ruszczyński, A · 2020
Later among the works it cites.
Distributionally robust policy evaluation and learning in offline contextual bandits
Si, N., Zhang, F., Zhou, Z., and Blanchet, J · 2020
Later among the works it cites.
Stable policy optimization via off-policy divergence regularization
Touati, A., Zhang, A., Pineau, J., and Vincent, P · 2020
Later among the works it cites.
Robust reinforcement learning using adversarial populations
Vinitsky, E., Du, Y., Parvate, K., Jang, K., Abbeel, P., and Bayen, A · 2020
Later among the works it cites.
Finite-sample analysis of Greedy-GQ with linear function approximation under Markovian noise
Wang, Y. and Zou, S · 2020
Later among the works it cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G · 2021
Later among the works it cites.
Robust reinforcement learning using least squares policy iteration with provable performance guarantees
Badrinath, K. P. and Kalathil, D · 2021
Later among the works it cites.
On the linear convergence of policy gradient methods for finite mdps
Bhandari, J. and Russo, D · 2021
Later among the works it cites.
Fast global convergence of natural policy gradient methods with entropy regularization
Cen, S., Cheng, C., Chen, Y., Wei, Y., and Chi, Y · 2021
Later among the works it cites.
Twice regularized MDPs and the equivalence between robustness and regularization
Derman, E., Geist, M., and Mannor, S · 2021
Later among the works it cites.
Maximum entropy RL (provably) solves some robust RL problems
Eysenbach, B. and Levine, S · 2021
Later among the works it cites.
Partial policy iteration for l1-robust markov decision processes
Ho, C. P., Petrik, M., and Wiesemann, W · 2021
Later among the works it cites.
Dr jekyll & mr hyde: the strange case of off-policy policy updates
Laroche, R. and des Combes, R. T · 2021
Later among the works it cites.
Softmax policy gradient methods can take exponential time to converge
Li, G., Wei, Y., Chi, Y., Gu, Y., and Chen, Y · 2021
Later among the works it cites.
Sample complexity of robust reinforcement learning with a generative model
Panaganti, K. and Kalathil, D · 2021
Later among the works it cites.
Online robust reinforcement learning with model uncertainty
Wang, Y. and Zou, S · 2021
Later among the works it cites.
Yang, W., Zhang, L., and Zhang, Z · 2021
Later among the works it cites.
Finite-sample regret bound for distributionally robust offline tabular reinforcement learning
Zhou, Z., Bai, Q., Zhou, Z., Qiu, L., Blanchet, J., and Glynn, P · 2021
Later among the works it cites.
On the convergence rates of policy gradient methods
Lin, X · 2022
Closest in time.