Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) studies how an agent comes to achieve reward in an environment through interactions over time.
On a measure of the information provided by an experiment
Lindley, D. V. et al · 1956
Earlier work this paper cites.
Expected information as expected utility
Bernardo, J. M · 1979
Earlier work this paper cites.
Toward a modern theory of adaptive networks: expectation and prediction
Sutton, R. S. & Barto, A. G · 1981
Earlier work this paper cites.
The role of theories in conceptual coherence
Murphy, G. L. & Medin, D. L · 1985
Earlier work this paper cites.
Conceptual Change in Childhood (MIT press, 1985)
Carey, S · 1985
Earlier work this paper cites.
Principles of object perception
Spelke, E. S · 1990
Earlier work this paper cites.
Q-learning
Watkins, C. J. & Dayan, P · 1992
Earlier work this paper cites.
A neural substrate of prediction and reward
Schultz, W., Dayan, P. & Montague, P. R · 1997
Earlier work this paper cites.
Words, thoughts, and theories , vol. 1 (Mit Press Cambridge, MA, 1997)
Gopnik, A., Meltzoff, A. N. & Bryant, P · 1997
Earlier work this paper cites.
Generalizing plans to new environments in relational mdps
Guestrin, C., Koller, D., Gearhart, C. & Kanodia, N · 2003
Earlier work this paper cites.
The essential child: Origins of essentialism in everyday thought (Oxford Series in Cognitive Dev, 2003)
Gelman, S. A · 2003
Earlier work this paper cites.
Infants’ physical world
Baillargeon, R · 2004
Earlier work this paper cites.
A theory of causal learning in children: causal maps and bayes nets
Gopnik, A. et al · 2004
Earlier work this paper cites.
Learning probabilistic relational planning rules
Pasula, H., Zettlemoyer, L. S. & Kaelbling, L. P · 2004
Earlier work this paper cites.
Cortical substrates for exploratory decisions in humans
Daw, N. D., O’doherty, J. P., Dayan, P., Seymour, B. & Dolan, R. J · 2006
Earlier work this paper cites.
God does not play dice: Causal determinism and preschoolers’ causal inferences
Schulz, L. E. & Sommerville, J · 2006
Earlier work this paper cites.
Core knowledge
Spelke, E. S. & Kinzler, K. D · 2007
Earlier work this paper cites.
Learning symbolic models of stochastic domains
Pasula, H. M., Zettlemoyer, L. S. & Kaelbling, L. P · 2007
Earlier work this paper cites.
Goal attribution to inanimate agents by 6.5-month-old infants
Csibra, G · 2008
Earlier work this paper cites.
Intuitive statistics by 8-month-old infants
Xu, F. & Garcia, V · 2008
Earlier work this paper cites.
Going beyond the evidence: Abstract laws and preschoolers’ responses to anomalous data
Schulz, L., Goodman, N. D., Tenenbaum, J. B. & Jenkins, A. C · 2008
Earlier work this paper cites.
An object-oriented representation for efficient reinforcement learning
Diuk, C., Cohen, A. & Littman, M. L · 2008
Earlier work this paper cites.
Theory-based causal induction
Griffiths, T. L. & Tenenbaum, J. B · 2009
Earlier work this paper cites.
The origin of concepts (Oxford University Press, 2009)
Carey, S · 2009
Earlier work this paper cites.
Learning to learn causal models
Kemp, C., Goodman, N. D. & Tenenbaum, J. B · 2010
Earlier work this paper cites.
Model-based influences on humans’ choices and striatal prediction errors
Daw, N. D., Gershman, S. J., Seymour, B., Dayan, P. & Dolan, R. J · 2011
Earlier work this paper cites.
Where science starts: Spontaneous experiments in preschoolers’ exploratory play
Cook, C., Goodman, N. D. & Schulz, L · 2011
Earlier work this paper cites.
Finding new facts; thinking new thoughts
Schulz, L · 2012
Earlier work this paper cites.
Width and serialization of classical planning problems (2012)
Geffner, H. & Lipovetzky, N · 2012
Cited alongside, same era.
The curse of planning: dissecting multiple reinforcement-learning systems by taxing the central executive
Otto, A. R., Gershman, S. J., Markman, A. B. & Daw, N. D · 2013
Cited alongside, same era.
A video game description language for model-based or interactive learning
Schaul, T · 2013
Cited alongside, same era.
Deep learning for real-time atari game play using offline monte-carlo tree search planning
Guo, X., Singh, S., Lee, H., Lewis, R. L. & Wang, X · 2014
Cited alongside, same era.
A physics-based model prior for object-oriented mdps
Scholz, J., Levihn, M., Isbell, C. & Wingate, D · 2014
Cited alongside, same era.
Information selection in noisy environments with large action spaces
Tsividis, P., Gershman, S., Tenenbaum, J. & Schulz, L · 2014
Best-first width search: Exploration and exploitation in classical planning
Lipovetzky, N. & Geffner, H · 2017
Later among the works it cites.
Avoiding frostbite: It helps to learn from others
Tessler, M. H., Goodman, N. D. & Frank, M. C · 2017
Later among the works it cites.
Evolution strategies as a scalable alternative to reinforcement learning
Salimans, T., Ho, J., Chen, X., Sidor, S. & Sutskever, I · 2017
Later among the works it cites.
Visual interaction networks: Learning a physics simulator from video
Watters, N. et al · 2017
Later among the works it cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A. & Darrell, T · 2017
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V. et al · 2015
Cited alongside, same era.
Schaul, T., Quan, J., Antonoglou, I. & Silver, D · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
Stadie, B. C., Levine, S. & Abbeel, P · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
Oh, J., Guo, X., Lee, H., Lewis, R. L. & Singh, S · 2015
Cited alongside, same era.
Learning visual predictive models of physics for playing billiards
Fragkiadaki, K., Agrawal, P., Levine, S. & Malik, J · 2015
Cited alongside, same era.
Hypothesis-space constraints in causal learning
Tsividis, P., Tenenbaum, J. B. & Schulz, L · 2015
Cited alongside, same era.
Silver, D. et al · 2018
Later among the works it cites.
Jaderberg, M. et al · 2018
Later among the works it cites.
Reinforcement learning: An introduction (MIT press, 2018)
Sutton, R. S. & Barto, A. G · 2018
Later among the works it cites.
Learning sparse relational transition models
Xia, V., Wang, Z. & Kaelbling, L. P · 2018
Later among the works it cites.
Strategic object oriented reinforcement learning
Keramati, R., Whang, J., Cho, P. & Brunskill, E · 2018
Later among the works it cites.
Ha, D. & Schmidhuber, J · 2018
Later among the works it cites.
Dreamcoder: Bootstrapping domain-specific languages for neurally-guided bayesian program learning
Ellis, K., Morales, L., Meyer, M. S., Solar-Lezama, A. & Tenenbaum, J. B · 2018
Later among the works it cites.
Prefrontal cortex as a meta-reinforcement learning system
Wang, J. X. et al · 2018
Later among the works it cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L. et al · 2018
Later among the works it cites.
Relational inductive biases, deep learning, and graph networks
Battaglia, P. W. et al · 2018
Later among the works it cites.
Relational deep reinforcement learning
Zambaldi, V. et al · 2018
Later among the works it cites.
Dopamine: A Research Framework for Deep Reinforcement Learning
Castro, P. S., Moitra, S., Gelada, C., Kumar, S. & Bellemare, M. G · 2018
Later among the works it cites.
Learning to play with intrinsically-motivated, self-aware agents
Haber, N., Mrowca, D., Wang, S., Fei-Fei, L. F. & Yamins, D. L · 2018
Later among the works it cites.
Alphastar: Mastering the real-time strategy game starcraft ii (2019)
Vinyals, O. et al · 2019
Later among the works it cites.
Recurrent experience replay in distributed reinforcement learning
Kapturowski, S., Ostrovski, G., Dabney, W., Quan, J. & Munos, R · 2019
Later among the works it cites.
Model-based reinforcement learning for atari
Kaiser, L. et al · 2019
Later among the works it cites.
Watters, N., Matthey, L., Bosnjak, M., Burgess, C. P. & Lerchner, A · 2019
Later among the works it cites.
Causal reasoning from meta-reinforcement learning
Dasgupta, I. et al · 2019
Later among the works it cites.
Agent57: Outperforming the Atari human benchmark
Badia, A. P. et al · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J. et al · 2020
Later among the works it cites.
Rapid trial-and-error learning with simulation supports flexible tool use and physical reasoning
Allen, K., Smith, K. & Tenenbaum, J. B · 2020
Later among the works it cites.
Mastering atari with discrete world models
Hafner, D., Lillicrap, T. P., Norouzi, M. & Ba, J · 2021
Closest in time.
What is the model in model-based planning?
Pouncy, T., Tsividis, P. & Gershman, S. J · 2021
Closest in time.