Fetching the paper…
Reading the bibliography…
Most of the recent deep reinforcement learning advances take an RL-centric perspective and focus on refinements of the training objective.
Eigenvalue computation in the 20th century
Golub, G. H. and van der Vorst, H. A · 2000
Earlier work this paper cites.
Deep learning via hessian-free optimization
Martens, J · 2010
Earlier work this paper cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Quasi newton temporal difference learning
Givchi, A. and Palhang, M · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Revisiting natural gradient for deep networks
Pascanu, R. and Bengio, Y · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
Martens, J. and Grosse, R · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Deep exploration via bootstrapped dqn
Osband, I., Blundell, C., Pritzel, A., and Roy, B. V · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, T. and Kingma, D. P · 2016
Earlier work this paper cites.
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., Hasselt, H., Lanctot, M., and Freitas, N · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning, 2017
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Earlier work this paper cites.
Parseval networks: Improving robustness to adversarial examples
Cissé, M., Bojanowski, P., Grave, E., Dauphin, Y., and Usunier, N · 2017
Earlier work this paper cites.
Improved training of wasserstein gans
Gulrajani, I., Ahmed, F., Arjovsky, M., Dumoulin, V., and Courville, A. C · 2017
Cited alongside, same era.
L2 regularization versus batch and weight normalization
Van Laarhoven, T · 2017
Cited alongside, same era.
Spectral norm regularization for improving the generalizability of deep learning
Yoshida, Y. and Miyato, T · 2017
Cited alongside, same era.
Theoretical analysis of auto rate-tuning by batch normalization
Arora, S., Li, Z., and Lyu, K · 2018
Cited alongside, same era.
Dopamine: A research framework for deep reinforcement learning
Castro, P. S., Moitra, S., Gelada, C., Kumar, S., and Bellemare, M. G · 2018
Quantifying generalization in reinforcement learning
Cobbe, K., Klimov, O., Hesse, C., Kim, T., and Schulman, J · 2019
Later among the works it cites.
Distilling policy distillation
Czarnecki, W. M., Pascanu, R., Osindero, S., Jayakumar, S., Swirszcz, G., and Jaderberg, M · 2019
Later among the works it cites.
Limitations of the empirical fisher approximation for natural gradient descent
Kunstner, F., Hennig, P., and Balles, L · 2019
Later among the works it cites.
A large-scale study on regularization and normalization in gans
Kurach, K., Lučić, M., Zhai, X., Michalski, M., and Gelly, S · 2019
Later among the works it cites.
Regularization matters in policy optimization – an empirical study on continuous control, 2019
Liu, Z., Li, X., Kang, B., and Darrell, T · 2019
Later among the works it cites.
Is deep reinforcement learning really superhuman on atari? leveling the playing field, 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Generalizable adversarial training via spectral normalization
Farnia, F., Zhang, J. M., and Tse, D · 2018
Cited alongside, same era.
Noisy networks for exploration
Fortunato, M., Azar, M. G., Piot, B., Menick, J., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., Pietquin, O., Blundell, C., and Legg, S · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Hasselt, H. V., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M. G., and Silver, D · 2018
Cited alongside, same era.
Norm matters: efficient and accurate normalization schemes in deep networks
Hoffer, E., Banner, R., Golan, I., and Soudry, D · 2018
Cited alongside, same era.
Limitations of the lipschitz constant as a defense against adversarial examples
Huster, T. P., Chiang, C. J., and Chadha, R · 2018
Cited alongside, same era.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M · 2018
Cited alongside, same era.
Spectral normalization for generative adversarial networks, 2018
Miyato, T., Kataoka, T., Koyama, M., and Yoshida, Y · 2018
Cited alongside, same era.
Toromanoff, M., Wirbel, E., and Moutarde, F · 2019
Later among the works it cites.
Minatar: An atari-inspired testbed for thorough and reproducible reinforcement learning experiments
Young, K. and Tian, T · 2019
Later among the works it cites.
Interference and generalization in temporal difference learning
Bengio, E., Pineau, J., and Precup, D · 2020
Later among the works it cites.
Regularisation of neural networks by enforcing lipschitz continuity
Gouk, H., Frank, E., Pfahringer, B., and Cree, M. J · 2020
Later among the works it cites.
Implicit under-parameterization inhibits data-efficient deep reinforcement learning
Kumar, A., Agarwal, R., Ghosh, D., and Levine, S · 2020
Later among the works it cites.
Simple and principled uncertainty estimation with deterministic deep learning via distance awareness
Liu, J. Z., Lin, Z., Padhy, S., Tran, D., Bedrax-Weiss, T., and Lakshminarayanan, B · 2020
Later among the works it cites.
Revisiting rainbow: Promoting more insightful and inclusive deep reinforcement learning research
Obando-Ceron, J. S. and Castro, P. S · 2020
Later among the works it cites.
Tdprop: Does jacobi preconditioning help temporal difference learning?
Romoff, J., Henderson, P., Kanaa, D., Bengio, E., Touati, A., Bacon, P., and Pineau, J · 2020
Later among the works it cites.
A case for new neural network smoothness constraints, 2020
Rosca, M., Weber, T., Gretton, A., and Mohamed, S · 2020
Later among the works it cites.
Adaptive temporal difference learning with linear function approximation
Sun, T., Shen, H., Chen, T., and Li, D · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization, 2020
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J., Levine, S., Finn, C., and Ma, T · 2020
Later among the works it cites.