Fetching the paper…
Reading the bibliography…
Evaluation of deep reinforcement learning (RL) is inherently challenging.
Natural environment benchmarks for reinforcement learning
Zhang, A., Wu, Y., and Pineau, J · 1905
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S., Barto, A. G., et al · 1998
Earlier work this paper cites.
Generalizing plans to new environments in relational mdps
Guestrin, C., Koller, D., Gearhart, C., and Kanodia, N · 2003
Earlier work this paper cites.
An object-oriented representation for efficient reinforcement learning
Diuk, C., Cohen, A., and Littman, M. L · 2008
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
Taylor, M. E. and Stone, P · 2009
Earlier work this paper cites.
Investigating contingency awareness using atari 2600 games
Bellemare, M. G., Veness, J., and Bowling, M · 2012
Earlier work this paper cites.
Fast reinforcement learning with large action sets using error-correcting output codes for mdp factorization
Dulac-Arnold, G., Denoyer, L., Preux, P., and Gallinari, P · 2012
Earlier work this paper cites.
The Arcade Learning Environment: An Evaluation Platform for General Agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps, 2013
Simonyan, K., Vedaldi, A., and Zisserman, A · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R · 2014
Earlier work this paper cites.
Frame skip is a powerful parameter for learning to play atari
Braylan, A., Hollenbeck, M., Meyerson, E., and Miikkulainen, R · 2015
Earlier work this paper cites.
Deep reinforcement learning in large discrete action spaces
Dulac-Arnold, G., Evans, R., van Hasselt, H., Sunehag, P., Lillicrap, T., Hunt, J., Mann, T., Weber, T., Degris, T., and Coppin, B · 2015
Earlier work this paper cites.
The dependence of effective planning horizon on model accuracy
Jiang, N., Kulesza, A., Singh, S., and Lewis, R · 2015
Earlier work this paper cites.
Human-level Control through Deep Reinforcement Learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Earlier work this paper cites.
Massively parallel methods for deep reinforcement learning
Nair, A., Srinivasan, P., Blackwell, S., Alcicek, C., Fearon, R., De Maria, A., Panneershelvam, V., Suleyman, M., Beattie, C., Petersen, S., et al · 2015
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Learning programs from noisy data
Raychev, V., Bielik, P., Vechev, M., and Krause, A · 2016
Cited alongside, same era.
Frame Skipping and Pre-Processing for Deep Q-Networks on Atari 2600 Games
Seita, D · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Cited alongside, same era.
OpenAI Baselines
Dhariwal, P., Hesse, C., Klimov, O., Nichol, A., Plappert, M., Radford, A., Schulman, J., Sidor, S., and Wu, Y · 2017
Cited alongside, same era.
Deep Reinforcement Learning that Matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2017
Cited alongside, same era.
Adversarial attacks on neural network policies
Huang, S., Papernot, N., Goodfellow, I., Duan, Y., and Abbeel, P · 2017
Cited alongside, same era.
Re-evaluating evaluation
Balduzzi, D., Tuyls, K., Perolat, J., and Graepel, T · 2018
Later among the works it cites.
Let’s Play Again: Variability of Deep Reinforcement Learning Agents in Atari Environments
Clary, K., Tosch, E., Foley, J., and Jensen, D · 2018
Later among the works it cites.
Quantifying generalization in reinforcement learning
Cobbe, K., Klimov, O., Hesse, C., Kim, T., and Schulman, J · 2018
Later among the works it cites.
Investigating human priors for playing video games
Dubey, R., Agrawal, P., Pathak, D., Griffiths, T. L., and Efros, A. A · 2018
Later among the works it cites.
Montezuma’s Revenge Solved by Go-Explore, a New Algorithm for Hard-Exploration Problems (Sets Records on Pitfall, Too)
Ecoffet, A., Huizinga, J., Lehman, J., Stanley, K. O., and Clune, J · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Schema networks: Zero-shot transfer with a generative causal model of intuitive physics
Kansky, K., Silver, T., Mély, D. A., Eldawy, M., Lázaro-Gredilla, M., Lou, X., Dorfman, N., Sidor, S., Phoenix, S., and George, D · 2017
Cited alongside, same era.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al · 2017
Cited alongside, same era.
Building machines that learn and think like people
Lake, B. M., Ullman, T. D., Tenenbaum, J. B., and Gershman, S. J · 2017
Cited alongside, same era.
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M. J., and Bowling, M · 2017
Cited alongside, same era.
Adversarially robust policy learning: Active construction of physically-plausible perturbations
Mandlekar, A., Zhu, Y., Garg, A., Fei-Fei, L., and Savarese, S · 2017
Cited alongside, same era.
Robust adversarial reinforcement learning
Pinto, L., Davidson, J., Sukthankar, R., and Gupta, A · 2017
Cited alongside, same era.
Treeqn and atreec: Differentiable tree-structured models for deep reinforcement learning
Farquhar, G., Rocktaschel, T., Igl, M., and Whiteson, S · 2018
Later among the works it cites.
Visualizing and understanding atari agents
Greydanus, S., Koul, A., Dodge, J., and Fern, A · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D · 2018
Later among the works it cites.
Deep reinforcement learning doesn’t work yet
Irpan, A · 2018
Later among the works it cites.
Using Cumulative Distribution Based Performance Analysis to Bechmark Models
Jordan, S. M., Cohen, D., and Thomas, P. S · 2018
Later among the works it cites.
Strategic object oriented reinforcement learning
Keramati, R., Whang, J., Cho, P., and Brunskill, E · 2018
Later among the works it cites.
Modularization of end-to-end learning: Case study in arcade games
Melnik, A., Fleer, S., Schilling, M., and Ritter, H · 2018
Later among the works it cites.
Gotta learn fast: A new benchmark for generalization in rl
Nichol, A., Pfau, V., Hesse, C., Klimov, O., and Schulman, J · 2018
Later among the works it cites.
Learning Montezuma’s Revenge from a Single Demonstration
Salimans, T. and Chen, R · 2018
Later among the works it cites.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O · 2019
Closest in time.
Wang, R., Lehman, J., Clune, J., and Stanley, K. O · 2019
Closest in time.