Fetching the paper…
Reading the bibliography…
The Arcade Learning Environment (ALE) is an evaluation platform that poses the challenge of building AI agents with general competency across dozens of Atari 2600 games.
Graying the Black Box: Understanding DQNs
Zahavy, T., Ben-Zrihem, N., and Mannor, S. (2016) · 1908
Earlier work this paper cites.
Asynchronous Methods for Deep Reinforcement Learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 1937
Earlier work this paper cites.
Dynamic Programming
Bellman, R. E. (1957) · 1957
Earlier work this paper cites.
Knowledge Growth in an Artificial Animal
Wilson, S. (1985) · 1985
Earlier work this paper cites.
Reinforcement Learning for Robots Using Neural Networks
Lin, L.-J. (1993) · 1993
Earlier work this paper cites.
Lifelong Robot Learning
Thrun, S., and Mitchell, T. M. (1993) · 1993
Earlier work this paper cites.
On-line Q-Learning using Connectionist Systems
Rummery, G. A., and Niranjan, M. (1994) · 1994
Earlier work this paper cites.
CHILD: A First Step Towards Continual Learning
Ring, M. (1997) · 1997
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S., and Barto, A. G. (1998) · 1998
Earlier work this paper cites.
Approximate Planning in Large POMDPs via Reusable Trajectories
Kearns, M. J., Mansour, Y., and Ng, A. Y. (1999) · 1999
Earlier work this paper cites.
R-MAX - A General Polynomial Time Algorithm for Near-Optimal Reinforcement Learning
Brafman, R. I., and Tennenholtz, M. (2002) · 2002
Earlier work this paper cites.
Deep Blue
Campbell, M., Jr., A. J. H., and Hsu, F. (2002) · 2002
Earlier work this paper cites.
Near-Optimal Reinforcement Learning in Polynomial Time
Kearns, M. J., and Singh, S. P. (2002) · 2002
Earlier work this paper cites.
Intrinsically Motivated Reinforcement Learning
Singh, S., Barto, A. G., and Chentanez, N. (2004) · 2004
Earlier work this paper cites.
Reinforcement Learning in POMDPs Without Resets
Even-Dar, E., Kakade, S. M., and Mansour, Y. (2005) · 2005
Earlier work this paper cites.
Universal Artificial Intelligence: Sequential Decisions based on Algorithmic Probability
Hutter, M. (2005) · 2005
Earlier work this paper cites.
Intrinsic Motivation Systems for Autonomous Mental Development
Oudeyer, P., Kaplan, F., and Hafner, V. (2007) · 2007
Earlier work this paper cites.
Checkers is Solved
Schaeffer, J., Burch, N., Björnsson, Y., Kishimoto, A., Müller, M., Lake, R., Lu, P., and Sutphen, S. (2007) · 2007
Earlier work this paper cites.
An Analysis of Model-Based Interval Estimation for Markov Decision Processes
Strehl, A. L., and Littman, M. L. (2008) · 2008
Earlier work this paper cites.
Racing the Beam: The Atari Video Computer System
Montfort, N., and Bogost, I. (2009) · 2009
Earlier work this paper cites.
Transfer Learning for Reinforcement Learning Domains: A Survey
Taylor, M. E., and Stone, P. (2009) · 2009
Earlier work this paper cites.
GQ( λ \lambda ): A General Gradient Algorithm for Temporal-Difference Prediction Learning with Eligibility Traces
Maei, H. R., and Sutton, R. S. (2010) · 2010
Earlier work this paper cites.
Game-independent AI Agents for Playing Atari 2600 Console Games
Naddaf, Y. (2010) · 2010
Earlier work this paper cites.
Horde: A Scalable Real-Time Architecture for Learning Knowledge from Unsupervised Sensorimotor Interaction
Sutton, R., Modayil, J., Delp, M., Degris, T., Pilarski, P., White, A., and Precup, D. (2011) · 2011
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012) · 2012
Earlier work this paper cites.
Intrinsic Motivation and Reinforcement Learning
Barto, A. G. (2013) · 2013
Cited alongside, same era.
Bayesian Learning of Recursively Factored Environments
Bellemare, M., Veness, J., and Bowling, M. (2013) · 2013
Cited alongside, same era.
The Arcade Learning Environment: An Evaluation Platform for General Agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013) · 2013
Cited alongside, same era.
General Video Game Playing
Levine, J., Congdon, C. B., Ebner, M., Kendall, G., Lucas, S. M., Miikkulainen, R., Schaul, T., and Thompson, T. (2013) · 2013
Cited alongside, same era.
A Survey of Real-Time Strategy Game AI Research and Competition in StarCraft
Ontanon, S., Synnaeve, G., Uriarte, A., Richoux, F., Churchill, D., and Preuss, M. (2013) · 2013
Cited alongside, same era.
Skip Context Tree Switching
Bellemare, M. G., Veness, J., and Talvitie, E. (2014) · 2014
Cited alongside, same era.
The Malmo Platform for Artificial Intelligence Experimentation
Johnson, M., Hofmann, K., Hutton, T., and Bignell, D. (2016) · 2016
Later among the works it cites.
State of the Art Control of Atari Games Using Shallow Reinforcement Learning
Liang, Y., Machado, M. C., Talvitie, E., and Bowling, M. H. (2016) · 2016
Later among the works it cites.
Safe and Efficient Off-Policy Reinforcement Learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M. G. (2016) · 2016
Later among the works it cites.
Deep Exploration via Bootstrapped DQN
Osband, I., Blundell, C., Pritzel, A., and Roy, B. V. (2016) · 2016
Later among the works it cites.
Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning
Parisotto, E., Ba, L. J., and Salakhutdinov, R. (2016) · 2016
Later among the works it cites.
Policy Distillation
Rusu, A. A., Colmenarejo, S. G., Gucehre, C., Desjardins, G., Kirkpatrick, J., Pascanu, R., Mnih, V., Kavukcuoglu, K., and Hadsell, R. (2016) · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A Comparison of Learning Algorithms on the Arcade Learning Environment
Defazio, A., and Graepel, T. (2014) · 2014
Cited alongside, same era.
Deep Learning for Real-Time Atari Game Play Using Offline Monte-Carlo Tree Search Planning
Guo, X., Singh, S., Lee, H., Lewis, R. L., and Wang, X. (2014) · 2014
Cited alongside, same era.
A Neuroevolution Approach to General Atari Game Playing
Hausknecht, M. J., Lehman, J., Miikkulainen, R., and Stone, P. (2014) · 2014
Cited alongside, same era.
Model Regularization for Stable Sample Rollouts
Talvitie, E. (2014) · 2014
Cited alongside, same era.
Reports from the 2015 AAAI Workshop Program
Albrecht, S. V., L., J. C., Buckeridge, D. L., Botea, A., Caragea, C., Chi, C., Damoulas, T., Dilkina, B. N., Eaton, E., Fazli, P., Ganzfried, S., Lindauer, M. T., Machado, M. C., Malitsky, Y., Marcus, G., Meijer, S., Rossi, F., Shaban-Nejad, A., Thiebaux, S., Veloso, M. M., Walsh, T., Wang, C., Zhang, J., and Zheng, Y. (2015) · 2015
Cited alongside, same era.
The Arcade Learning Environment: An Evaluation Platform for General Agents (Extended Abstract)
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2015) · 2015
Cited alongside, same era.
Later among the works it cites.
Prioritized Experience Replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D. (2016) · 2016
Later among the works it cites.
Blind Search for Atari-Like Online Planning Revisited
Shleyfman, A., Tuisov, A., and Domshlak, C. (2016) · 2016
Later among the works it cites.
Mastering the Game of Go with Deep Neural Networks and Tree Search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D. (2016) · 2016
Later among the works it cites.
Dueling Network Architectures for Deep Reinforcement Learning
Wang, Z., Schaul, T., Hessel, M., van Hasselt, H., Lanctot, M., and De Freitas, N. (2016) · 2016
Later among the works it cites.
A Distributional Perspective on Reinforcement Learning
Bellemare, M. G., Dabney, W., and Munos, R. (2017) · 2017
Closest in time.
Recurrent Environment Simulators
Chiappa, S., Racaniere, S., Wierstra, D., and Mohamed, S. (2017) · 2017
Closest in time.
Value-Aware Loss Function for Model-based Reinforcement Learning
Farahmand, A.-M., Barreto, A., and Nikovski, D. (2017) · 2017
Closest in time.
Learning from Demonstrations for Real World Reinforcement Learning
Hester, T., Vecerik, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Sendonaris, A., Dulac-Arnold, G., Osband, I., Agapiou, J., et al. (2017) · 2017
Closest in time.
Reproducibility of Benchmarked Deep Reinforcement Learning Tasks for Continuous Control
Islam, R., Henderson, P., Gomrokchi, M., and Precup, D. (2017) · 2017
Closest in time.
Reinforcement Learning with Unsupervised Auxiliary Tasks
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K. (2017) · 2017
Closest in time.
Learning to Prune Dominated Action Sequences in Online Black-Box Planning
Jinnai, Y., and Fukunaga, A. (2017) · 2017
Closest in time.
Emergent Tangled Graph Representations for Atari Game Playing Agents
Kelly, S., and Heywood, M. I. (2017) · 2017
Closest in time.
A Laplacian Framework for Option Discovery in Reinforcement Learning
Machado, M. C., Bellemare, M. G., and Bowling, M. (2017) · 2017
Closest in time.
Count-Based Exploration in Feature Space for Reinforcement Learning
Martin, J., Sasikumar, S. N., Everitt, T., and Hutter, M. (2017) · 2017
Closest in time.
Count-Based Exploration with Neural Density Models
Ostrovski, G., Bellemare, M. G., van den Oord, A., and Munos, R. (2017) · 2017
Closest in time.
Self-correcting models for model-based reinforcement learning
Talvitie, E. (2017) · 2017
Closest in time.
FeUdal Networks for Hierarchical Reinforcement Learning
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K. (2017) · 2017
Closest in time.
Deep Reinforcement Learning with Double Q-Learning
van Hasselt, H., Guez, A., and Silver, D. (2016) · 2094
Closest in time.