Fetching the paper…
Reading the bibliography…
Modern policy gradient algorithms such as Proximal Policy Optimization (PPO) rely on an arsenal of heuristics, including loss clipping and gradient clipping, to ensure successful learning.
Why gradient clipping accelerates training: A theoretical justification for adaptivity
Zhang, J., He, T., Sra, S., and Jadbabaie, A · 1905
Earlier work this paper cites.
Why adam beats sgd for attention models
Zhang, J., Karimireddy, S. P., Veit, A., Kim, S., Reddi, S. J., Kumar, S., and Sra, S · 1912
Earlier work this paper cites.
A test of goodness of fit
Anderson, T. W. and Darling, D. A · 1954
Earlier work this paper cites.
A simple general approach to inference about the tail of a distribution
Hill, B. M · 1975
Earlier work this paper cites.
Problem Complexity and Method Efficiency in Optimization
Nemirovski, A. and Yudin, D · 1983
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M · 2002
Earlier work this paper cites.
Heavy-tail phenomena: probabilistic and statistical modeling
Resnick, S. I · 2007
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Bandits with heavy tail
Bubeck, S., Cesa-Bianchi, N., and Lugosi, G · 2013
Earlier work this paper cites.
Geometric median and robust estimation in banach spaces
Minsker, S. et al · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
On estimating the tail index and the spectral measure of multivariate α \alpha -stable distributions
Mohammadi, M., Mohammadpour, A., and Ogata, H · 2015
Cited alongside, same era.
Tail index estimation: Quantile driven threshold selection
Danielsson, J., Ergun, L. M., de Haan, L., and de Vries, C. G · 2016
Cited alongside, same era.
Q-prop: Sample-efficient policy gradient with an off-policy critic
Gu, S., Lillicrap, T., Ghahramani, Z., Turner, R. E., and Levine, S · 2016
Cited alongside, same era.
No-regret algorithms for heavy-tailed linear bandits
Medina, A. M. and Yang, S · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Robust estimation via robust gradient estimation
Prasad, A., Suggala, A. S., Balakrishnan, S., and Ravikumar, P · 2018
Later among the works it cites.
Variational inference with tail-adaptive f-divergence
Wang, D., Liu, H., and Liu, Q · 2018
Later among the works it cites.
Variance reduction for policy gradient with action-dependent factorized baselines
Wu, C., Rajeswaran, A., Duan, Y., Kumar, V., Bayen, A. M., Kakade, S., Mordatch, I., and Abbeel, P · 2018
Later among the works it cites.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Dkebiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al · 2019
Later among the works it cites.
Implementation matters in deep rl: A case study on ppo and trpo
Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., and Madry, A · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2017
Cited alongside, same era.
Reproducibility of benchmarked deep reinforcement learning tasks for continuous control
Islam, R., Henderson, P., Gomrokchi, M., and Precup, D · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Cited alongside, same era.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al · 2018
Cited alongside, same era.
Henderson, P., Romoff, J., and Pineau, J · 2018
Cited alongside, same era.
A closer look at deep policy gradients
Ilyas, A., Engstrom, L., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., and Madry, A · 2018
Cited alongside, same era.
Later among the works it cites.
Non-gaussianity of stochastic gradient noise
Panigrahi, A., Somani, R., Goyal, N., and Netrapalli, P · 2019
Later among the works it cites.
A tail-index analysis of stochastic gradient noise in deep neural networks
Simsekli, U., Sagun, L., and Gurbuzbalaban, M · 2019
Later among the works it cites.
Trajectory-wise control variates for variance reduction in policy gradient methods
Cheng, C.-A., Yan, X., and Boots, B · 2020
Later among the works it cites.
Beyond variance reduction: Understanding the true impact of baselines on policy optimization
Chung, W., Thomas, V., Machado, M. C., and Roux, N. L · 2020
Later among the works it cites.
Importance sampling techniques for policy optimization
Metelli, A. M., Papini, M., Montali, N., and Restelli, M · 2020
Later among the works it cites.
Şimşekli, U., Zhu, L., Teh, Y. W., and Gürbüzbalaban, M · 2020
Later among the works it cites.
A diffusion theory for deep learning dynamics: Stochastic gradient descent exponentially favors flat minima
Xie, Z., Sato, I., and Sugiyama, M · 2021
Closest in time.