Fetching the paper…
Reading the bibliography…
Off-policy reinforcement learning (RL) from pixel observations is notoriously unstable.
The probable error of a mean
Student · 1908
Earlier work this paper cites.
Individual comparisons by ranking methods
Wilcoxon, F · 1945
Earlier work this paper cites.
On a Test of Whether one of Two Random Variables is Stochastically Larger than the Other
Mann, H. B. and Whitney, D. R · 1947
Earlier work this paper cites.
A markovian decision process
Bellman, R · 1957
Earlier work this paper cites.
Learning to predict by the method of temporal differences
Sutton, R · 1988
Earlier work this paper cites.
Bootstrap methods: another look at the jackknife
Efron, B · 1992
Earlier work this paper cites.
On-line Q-learning using connectionist systems
Rummery, G. A. and Niranjan, M · 1994
Earlier work this paper cites.
Gradient descent for general reinforcement learning
Baird, L. and Moore, A · 1998
Earlier work this paper cites.
Reinforcement learning through gradient descent
Baird, L · 1999
Earlier work this paper cites.
Bias-variance error bounds for temporal difference updates
Kearns, M. J. and Singh, S. P · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
Benchmarking optimization software with performance profiles
Dolan, E. D. and Moré, J. J · 2002
Earlier work this paper cites.
Convex optimization
Boyd, S., Boyd, S. P., and Vandenberghe, L · 2004
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Learning visual feature spaces for robotic manipulation with deep spatial autoencoders
Finn, C., Tan, X. Y., Duan, Y., Darrell, T., Levine, S., and Abbeel, P · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Benchmarking deep reinforcement learning for continuous control
Duan, Y., Chen, X., Houthooft, R., Schulman, J., and Abbeel, P · 2016
Earlier work this paper cites.
A closer look at memorization in deep networks
Arpit, D., Jastrzebski, S., Ballas, N., Krueger, D., Bengio, E., Kanwal, M. S., Maharaj, T., Fischer, A., Courville, A., Bengio, Y., and Lacoste-Julien, S · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2017
Earlier work this paper cites.
Exploring generalization in deep learning
Neyshabur, B., Bhojanapalli, S., McAllester, D., and Srebro, N · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al · 2017
Cited alongside, same era.
Reinforcement learning for control: Performance, stability, and deep approximators
Buşoniu, L., de Bruin, T., Tolić, D., Kober, J., and Palunko, I · 2018
Cited alongside, same era.
Learning actionable representations from visual observations
Dwibedi, D., Tompson, J., Lynch, C., and Sermanet, P · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Fujimoto, S., van Hoof, H., and Meger, D · 2018
Cited alongside, same era.
Dropblock: A regularization method for convolutional networks
Ghiasi, G., Lin, T.-Y., and Le, Q. V · 2018
Cited alongside, same era.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M · 2020
Later among the works it cites.
Model based reinforcement learning for Atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osiński, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., Mohiuddin, A., Sepassi, R., Tucker, G., and Michalewski, H · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems, 2020
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Later among the works it cites.
Bad global minima exist and sgd can reach them
Liu, S., Papailiopoulos, D., and Achlioptas, D · 2020
Later among the works it cites.
What do neural networks learn when trained with random labels?
Maennel, H., Alabdulmohsin, I. M., Tolstikhin, I. O., Baldock, R., Bousquet, O., Gelly, S., and Keysers, D · 2020
Later among the works it cites.
A case for new neural networks smoothness constraints
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D · 2018
Cited alongside, same era.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., et al · 2018
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Li, H., Xu, Z., Taylor, G., Studer, C., and Goldstein, T · 2018
Cited alongside, same era.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M · 2018
Cited alongside, same era.
Spectral normalization for generative adversarial networks
Miyato, T., Kataoka, T., Koyama, M., and Yoshida, Y · 2018
Cited alongside, same era.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al · 2018
Cited alongside, same era.
Deep reinforcement learning and the deadly triad
Van Hasselt, H., Doron, Y., Strub, F., Hessel, M., Sonnerat, N., and Modayil, J · 2018
Cited alongside, same era.
Rosca, M., Weber, T., Gretton, A., and Mohamed, S · 2020
Later among the works it cites.
Data-efficient reinforcement learning with self-predictive representations
Schwarzer, M., Anand, A., Goel, R., Hjelm, R. D., Courville, A., and Bachman, P · 2020
Later among the works it cites.
Observational overfitting in reinforcement learning
Song, X., Jiang, Y., Tu, S., Du, Y., and Neyshabur, B · 2020
Later among the works it cites.
The ingredients of real-world robotic reinforcement learning
Zhu, H., Yu, J., Gupta, A., Shah, D., Hartikainen, K., Singh, A., Kumar, V., and Levine, S · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice, 2021
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A., and Bellemare, M. G · 2021
Later among the works it cites.
Feature purification: How adversarial training performs robust deep learning, 2021
Allen-Zhu, Z. and Li, Y · 2021
Later among the works it cites.
Mind the pad – {cnn}s can develop blind spots
Alsallakh, B., Kokhlikyan, N., Miglani, V., Yuan, J., and Reblitz-Richardson, O · 2021
Later among the works it cites.
Offline RL without off-policy evaluation
Brandfonbrener, D., Whitney, W. F., Ranganath, R., and Bruna, J · 2021
Later among the works it cites.
Learning pessimism for robust and efficient off-policy reinforcement learning
Cetin, E. and Celiktutan, O · 2021
Later among the works it cites.
Phasic policy gradient
Cobbe, K. W., Hilton, J., Klimov, O., and Schulman, J · 2021
Later among the works it cites.
D4{rl}: Datasets for deep data-driven reinforcement learning, 2021
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2021
Later among the works it cites.
Spectral normalisation for deep reinforcement learning: An optimisation perspective
Gogianu, F., Berariu, T., Rosca, M. C., Clopath, C., Busoniu, L., and Pascanu, R · 2021
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Kostrikov, I., Yarats, D., and Fergus, R · 2021
Later among the works it cites.
Implicit under-parameterization inhibits data-efficient deep reinforcement learning
Kumar, A., Agarwal, R., Ghosh, D., and Levine, S · 2021
Later among the works it cites.
Tactical optimism and pessimism for deep reinforcement learning
Moskovitz, T., Parker-Holder, J., Pacchiano, A., Arbel, M., and Jordan, M · 2021
Later among the works it cites.
Return-based scaling: Yet another normalisation trick for deep rl, 2021
Schaul, T., Ostrovski, G., Kemaev, I., and Borsa, D · 2021
Later among the works it cites.
Improving sample efficiency in model-free reinforcement learning from images
Yarats, D., Zhang, A., Kostrikov, I., Amos, B., Pineau, J., and Fergus, R · 2021
Later among the works it cites.
Automated reinforcement learning (autorl): A survey and open problems, 2022
Parker-Holder, J., Rajan, R., Song, X., Biedenkapp, A., Miao, Y., Eimer, T., Zhang, B., Nguyen, V., Calandra, R., Faust, A., Hutter, F., and Lindauer, M · 2022
Closest in time.
Mastering visual continuous control: Improved data-augmented reinforcement learning
Yarats, D., Fergus, R., Lazaric, A., and Pinto, L · 2022
Closest in time.