Fetching the paper…
Reading the bibliography…
We focus on parameterized policy search for reinforcement learning over continuous action spaces.
Dynamic Programming
Bellman, R. E · 1957
Earlier work this paper cites.
Monotone operators and the proximal point algorithm
Rockafellar, R. T · 1976
Earlier work this paper cites.
Fractals and self similarity
Hutchinson, J. E · 1981
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Actor-critic–type learning algorithms for Markov decision processes
Konda, V. R. and Borkar, V. S · 1999
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Baxter, J. and Bartlett, P. L · 2001
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M · 2002
Earlier work this paper cites.
Emergence of scaling in complex networks
Barabási, A.-L. et al · 2003
Earlier work this paper cites.
Learning rates for q-learning
Even-Dar, E., Mansour, Y., and Bartlett, P · 2003
Earlier work this paper cites.
Fat tails, scaling, and stable laws: a critical look at modeling extremal events in financial phenomena
Focardi, S. M. and Fabozzi, F. J · 2003
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications
Kushner, H. J. and Yin, G. G · 2003
Earlier work this paper cites.
Stochastic optimal control: the discrete-time case
Bertsekas, D. P. and Shreve, S · 2004
Earlier work this paper cites.
Stochastic approximation: A dynamical systems viewpoint
Borkar, V. S · 2008
Earlier work this paper cites.
Natural actor-critic algorithms
Bhatnagar, S., Sutton, R., Ghavamzadeh, M., and Lee, M · 2009
Cited alongside, same era.
Robust stochastic approximation approach to stochastic programming
Nemirovski, A., Juditsky, A., Lan, G., and Shapiro, A · 2009
Cited alongside, same era.
Sample efficient reinforcement learning with reinforce
Zhang, J., Kim, J., O’Donoghue, B., and Boyd, S · 2010
Cited alongside, same era.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Ghadimi, S. and Lan, G · 2013
Cited alongside, same era.
Markov Decision Processes: Discrete stochastic dynamic programming
Puterman, M. L · 2014
Cited alongside, same era.
Adaptive treatment strategies in practice: planning trials and analyzing data for personalized medicine
Policy gradient using weak derivatives for reinforcement learning
Bhatt, S., Koppel, A., and Krishnamurthy, V · 2019
Later among the works it cites.
Momentum-based variance reduction in non-convex sgd
Cutkosky, A. and Orabona, F · 2019
Later among the works it cites.
The problem with ddpg: understanding failures in deterministic environments with sparse rewards
Matheron, G., Perrin, N., and Sigaud, O · 2019
Later among the works it cites.
First exit time analysis of stochastic gradient descent under heavy-tailed gradient noise
Nguyen, T. H., 𝒮 \mathcal{S} im 𝓈 \mathcal{s} ekli, U., Gürbüzbalaban, M., and Richard, G · 2019
Later among the works it cites.
Sample efficient policy gradient methods with recursive variance reduction
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kosorok, M. R. and Moodie, E. E · 2015
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P · 2015
Cited alongside, same era.
Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
Ghadimi, S., Lan, G., and Zhang, H · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Cited alongside, same era.
Improving stochastic policy gradients in continuous control with deep reinforcement learning using the beta distribution
Chou, P.-W., Maturana, D., and Scherer, S · 2017
Cited alongside, same era.
Zap q-learning
Devraj, A. M. and Meyn, S. P · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Xu, P., Gao, F., and Gu, Q · 2019
Later among the works it cites.
Policy optimization with stochastic mirror descent
Yang, L., Zheng, G., Zhang, H., Zhang, Y., Zheng, Q., Wen, J., and Pan, G · 2019
Later among the works it cites.
Reinforcement learning to optimize long-term user engagement in recommender systems
Zou, L., Xia, L., Ding, Z., Song, J., Liu, W., and Yin, D · 2019
Later among the works it cites.
Optimality and approximation with policy gradient methods in markov decision processes
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G · 2020
Later among the works it cites.
Balancing learning speed and stability in policy gradient via adaptive exploration
Papini, M., Battistello, A., and Restelli, M · 2020
Later among the works it cites.
Mirror descent policy optimization
Tomar, M., Shani, L., Efroni, Y., and Ghavamzadeh, M · 2020
Later among the works it cites.
Stochastic recursive momentum for policy gradient methods
Yuan, H., Lian, X., Liu, J., and Zhou, Y · 2020
Later among the works it cites.
On the sample complexity and metastability of heavy-tailed policy search in continuous control
Bedi, A. S., Parayil, A., Zhang, J., Wang, M., and Koppel, A · 2021
Later among the works it cites.
On proximal policy optimization’s heavy-tailed gradients
Garg, S., Zhanson, J., Parisotto, E., Prasad, A., Kolter, J. Z., Balakrishnan, S., Lipton, Z. C., Salakhutdinov, R., and Ravikumar, P · 2021
Later among the works it cites.
On the linear convergence of natural policy gradient algorithm
Khodadadian, S., Jhunjhunwala, P. R., Varma, S. M., and Maguluri, S. T · 2021
Later among the works it cites.
Lan, G · 2021
Later among the works it cites.