Fetching the paper…
Reading the bibliography…
In the context of average-reward reinforcement learning, the requirement for oracle knowledge of the mixing time, a measure of the duration a Markov chain under a fixed policy needs to achieve its stationary distribution, poses a significant challenge for the global convergence of policy gradient methods.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D., Singh, S., and Mansour, Y · 1999
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Ergodic mirror descent
Duchi, J. C., Agarwal, A., Johansson, M., and Jordan, M. I · 2012
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
Mixing time estimation in reversible markov chains from a single sample path
Hsu, D. J., Kontorovich, A., and Szepesvári, C · 2015
Earlier work this paper cites.
Online to offline conversions, universality and adaptive minibatch sizes
Levy, K · 2017
Earlier work this paper cites.
Stochastic variance-reduced policy gradient
Papini, M., Binaghi, D., Canonaco, G., Pirotta, M., and Restelli, M · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
Kumar, H., Koppel, A., and Ribeiro, A · 2019
Earlier work this paper cites.
Neural policy gradient methods: Global optimality and rates of convergence
Wang, L., Cai, Q., Yang, Z., and Wang, Z · 2019
Earlier work this paper cites.
Optimality and approximation with policy gradient methods in markov decision processes
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G · 2020
Cited alongside, same era.
A multi-agent reinforcement learning perspective on distributed traffic engineering
Geng, N., Lan, T., Aggarwal, V., Yang, Y., and Xu, M · 2020
Cited alongside, same era.
A duality approach for regret minimization in average-award ergodic markov decision processes
Gong, H. and Wang, M · 2020
Cited alongside, same era.
An improved analysis of (variance-reduced) policy gradient and natural policy gradient methods
Liu, Y., Zhang, K., Basar, T., and Yin, W · 2020
Cited alongside, same era.
On the global convergence rates of softmax policy gradient methods
Mei, J., Xiao, C., Szepesvari, C., and Schuurmans, D · 2020
Cited alongside, same era.
Least squares regression with markovian data: Fundamental limits and algorithms
Learning infinite-horizon average-reward mdps with linear function approximation
Wei, C.-Y., Jahromi, M. J., Luo, H., and Jain, R · 2021
Later among the works it cites.
On-policy deep reinforcement learning for the average-reward criterion
Zhang, Y. and Ross, K. W · 2021
Later among the works it cites.
On the hidden biases of policy mirror ascent in continuous action spaces
Bedi, A. S., Chakraborty, S., Parayil, A., Sadler, B. M., Tokekar, P., and Koppel, A · 2022
Later among the works it cites.
Adapting to mixing time in stochastic optimization with Markovian data
Dorfman, R. and Levy, K. Y · 2022
Later among the works it cites.
Imed-rl: Regret optimal learning of ergodic markov decision processes
Pesquerel, F. and Maillard, O.-A · 2022
Later among the works it cites.
Cooperating graph neural networks with deep reinforcement learning for vaccine prioritization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nagaraj, D., Wu, X., Bresler, G., Jain, P., and Netrapalli, P · 2020
Cited alongside, same era.
Mixing time estimation in ergodic markov chains from a single trajectory with contraction methods
Wolfer, G · 2020
Cited alongside, same era.
An improved convergence analysis of stochastic variance-reduced policy gradient
Xu, P., Gao, F., and Gu, Q · 2020
Cited alongside, same era.
Global convergence of policy gradient methods to (almost) locally optimal policies
Zhang, K., Koppel, A., Zhu, H., and Basar, T · 2020
Cited alongside, same era.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G · 2021
Cited alongside, same era.
Continual learning in environments with polynomial mixing times
Riemer, M., Raparthy, S. C., Cases, I., Subbaraj, G., Touzel, M. P., and Rish, I · 2021
Cited alongside, same era.
Ling, L., Mondal, W. U., and Ukkusuri, S. V · 2023
Later among the works it cites.
Ada-nav: Adaptive trajectory-based sample efficient policy learning for robotic navigation
Patel, B., Weerakoon, K., Suttle, W. A., Koppel, A., Sadler, B. M., Bedi, A. S., and Manocha, D · 2023
Later among the works it cites.
Beyond exponentially fast mixing in average-reward reinforcement learning via multi-level monte carlo actor-critic
Suttle, W. A., Bedi, A., Patel, B., Sadler, B. M., Koppel, A., and Manocha, D · 2023
Later among the works it cites.
Regret analysis of policy gradient algorithm for infinite horizon average reward markov decision processes
Bai, Q., Mondal, W. U., and Aggarwal, V · 2024
Closest in time.
Global optimality guarantees for policy gradient methods
Bhandari, J. and Russo, D · 2024
Closest in time.