Fetching the paper…
Reading the bibliography…
In this article we introduce the Arcade Learning Environment (ALE): both a challenge problem and a platform and methodology for evaluating the development of general, domain-independent AI technology.
Generalized polynomial approximations in Markovian decision processes
Schweitzer, P. J., and Seidmann, A. (1985) · 1985
Earlier work this paper cites.
Sparse Distributed Memory
Kanerva, P. (1988) · 1988
Earlier work this paper cites.
Q-learning
Watkins, C., and Dayan, P. (1992) · 1992
Earlier work this paper cites.
Strategy Generation and Evaluation for Meta-Game Playing
Pell, B. (1993) · 1993
Earlier work this paper cites.
Lifelong robot learning
Thrun, S., and Mitchell, T. M. (1995) · 1995
Earlier work this paper cites.
Map learning with uninterpreted sensors and effectors
Pierce, D., and Kuipers, B. (1997) · 1997
Earlier work this paper cites.
Rationality and intelligence
Russell, S. J. (1997) · 1997
Earlier work this paper cites.
A non-behavioural, computational extension to the Turing Test
Dowe, D. L., and Hajek, A. R. (1998) · 1998
Earlier work this paper cites.
A formal definition of intelligence based on an intensional variant of Kolmogorov complexity
Hernández-Orallo, J., and Minaya-Collado, N. (1998) · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S., and Barto, A. G. (1998) · 1998
Earlier work this paper cites.
Similarity search in high dimensions via hashing
Gionis, A., Indyk, P., and Motwani, R. (1999) · 1999
Earlier work this paper cites.
General Game Playing: Overview of the AAAI competition
Genesereth, M. R., Love, N., and Pell, B. (2005) · 2005
Cited alongside, same era.
Universal Artificial Intelligence: Sequential Decisions based on Algorithmic Probability
Hutter, M. (2005) · 2005
Cited alongside, same era.
Bandit based Monte-Carlo planning
Kocsis, L., and Szepesvári, C. (2006) · 2006
Cited alongside, same era.
Coevolution of neural networks using a layered pareto archive
Monroy, G. A., Stanley, K. O., and Miikkulainen, R. (2006) · 2006
Cited alongside, same era.
An object-oriented representation for efficient reinforcement learning
Diuk, C., Cohen, A., and Littman, M. L. (2008) · 2008
Cited alongside, same era.
Machine Super Intelligence
Legg, S. (2008) · 2008
Cited alongside, same era.
The reinforcement learning competitions
Whiteson, S., Tanner, B., and White, A. (2010) · 2010
Later among the works it cites.
Using imagery to simplify perceptual abstraction in reinforcement learning agents
Wintermute, S. (2010) · 2010
Later among the works it cites.
Automatic state abstraction from demonstration
Cobo, L. C., Zang, P., Isbell, C. L., and Thomaz, A. L. (2011) · 2011
Later among the works it cites.
An approximation of the universal intelligence measure
Legg, S., and Veness, J. (2011) · 2011
Later among the works it cites.
Measuring intelligence through games
Schaul, T., Togelius, J., and Schmidhuber, J. (2011) · 2011
Later among the works it cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
From pixels to policies: A bootstrapping agent
Stober, J., and Kuipers, B. (2008) · 2008
Cited alongside, same era.
Learning to play Mario
Mohan, S., and Laird, J. E. (2009) · 2009
Cited alongside, same era.
Racing the Beam: The Atari Video Computer System
Montfort, N., and Bogost, I. (2009) · 2009
Cited alongside, same era.
Measuring universal intelligence: Towards an anytime intelligence test
Hernández-Orallo, J., and Dowe, D. L. (2010) · 2010
Cited alongside, same era.
Game-Independent AI Agents for Playing Atari 2600 Console Games
Naddaf, Y. (2010) · 2010
Cited alongside, same era.
Sutton, R., Modayil, J., Delp, M., Degris, T., Pilarski, P., White, A., and Precup, D. (2011) · 2011
Later among the works it cites.
Protecting against evaluation overfitting in empirical reinforcement learning
Whiteson, S., Tanner, B., Taylor, M. E., and Stone, P. (2011) · 2011
Later among the works it cites.
Investigating contingency awareness using Atari 2600 games
Bellemare, M., Veness, J., and Bowling, M. (2012) · 2012
Closest in time.
A survey of Monte Carlo tree search methods
Browne, C. B., Powley, E., Whitehouse, D., Lucas, S. M., Cowling, P. I., Rohlfshagen, P., Tavener, S., Perez, D., Samothrakis, S., and Colton, S. (2012) · 2012
Closest in time.
A survey of the seventh international planning competition
Coles, A., Coles, A., Olaya, A., Jiménez, S., López, C., Sanner, S., and Yoon, S. (2012) · 2012
Closest in time.
HyperNEAT-GGP: A HyperNEAT-based Atari general game player
Hausknecht, M., Khandelwal, P., Miikkulainen, R., and Stone, P. (2012) · 2012
Closest in time.