Fetching the paper…
Reading the bibliography…
We present an end-to-end, model-based deep reinforcement learning agent which dynamically attends to relevant parts of its state during planning.
Differentiable subset sampling
S. M. Xie and S. Ermon · 1901
Earlier work this paper cites.
Model-based reinforcement learning for atari
L. Kaiser, M. Babaeizadeh, P. Milos, B. Osinski, R. H. Campbell, K. Czechowski, D. Erhan, C. Finn, P. Kozakowski, S. Levine, et al · 1903
Earlier work this paper cites.
When to trust your model: Model-based policy optimization
M. Janner, J. Fu, M. Zhang, and S. Levine · 1906
Earlier work this paper cites.
Mastering Atari, Go, Chess and Shogi by planning with a learned model
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, et al · 1911
Earlier work this paper cites.
Every good regulator of a system must be a model of that system
R. C. Conant and W. Ross Ashby · 1970
Earlier work this paper cites.
Parametric correspondence and Chamfer matching: 2 new techniques for image matching
H. G. Barrow, J. M. Tenenbaum, R. C. Bolles, and H. C. Wolf · 1977
Earlier work this paper cites.
Hierarchical chamfer matching: A parametric edge matching algorithm
G. Borgefors · 1988
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
R. S. Sutton · 1991
Earlier work this paper cites.
A cognitive theory of consciousness
B. J. Baars · 1993
Earlier work this paper cites.
The conscious access hypothesis: origins and recent evidence
B. J. Baars · 2002
Earlier work this paper cites.
Investigating simple object representations in model-free deep reinforcement learning
G. Davidson and B. M. Lake · 2002
Earlier work this paper cites.
Consciousness
R. van Gulick · 2004
Earlier work this paper cites.
End-to-end object detection with transformers
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko · 2005
Earlier work this paper cites.
Robust constrained model predictive control
A. G. Richards · 2005
Earlier work this paper cites.
Conditional set generation with transformers
A. R. Kosiorek, H. Kim, and D. J. Rezende · 2006
Earlier work this paper cites.
Model-based reinforcement learning: A survey
T. M. Moerland, J. Broekens, and C. M. Jonker · 2006
Earlier work this paper cites.
D. Y.-T. Hui, M. Chevalier-Boisvert, D. Bahdanau, and Y. Bengio · 2007
Earlier work this paper cites.
Simple local models for complex dynamical systems
E. Talvitie and S. Singh · 2008
Cited alongside, same era.
A survey of numerical methods for optimal control
A. V. Rao · 2009
Cited alongside, same era.
Mastering Atari with discrete world models
D. Hafner, T. P. Lillicrap, M. Norouzi, and J. Ba · 2010
Cited alongside, same era.
Episodic memory for learning subjective-timescale models
A. Zakharov, M. Crosby, and Z. Fountas · 2010
Cited alongside, same era.
Inductive biases for deep learning of higher-level cognition
A. Goyal and Y. Bengio · 2011
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
M. Hessel, J. Modayil, H. van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. G. Azar, and D. Silver · 2017
Later among the works it cites.
Reinforcement learning with unsupervised auxiliary tasks
M. Jaderberg, V. Mnih, W. M. Czarnecki, T. Schaul, J. Z. Leibo, D. Silver, and K. Kavukcuoglu · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Later among the works it cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. B. Hamrick, A. L. Friesen, F. Behbahani, A. Guez, F. Viola, S. Witherspoon, T. Anthony, L. Buesing, P. Velickovic, and T. Weber · 2011
Cited alongside, same era.
Learning object-centric video models by contrasting sets
S. Löwe, K. Greff, R. Jonschkowski, A. Dosovitskiy, and T. Kipf · 2011
Cited alongside, same era.
Refactoring policy for compositional generalizability using self-supervised object proposals
T. Mu, J. Gu, Z. Jia, H. Tang, and H. Su · 2011
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Deep reinforcement learning with double q-learning
H. van Hasselt, A. Guez, and D. Silver · 2015
Cited alongside, same era.
M. Zaheer, S. Kottur, S. Ravanbhakhsh, B. Póczos, R. Salakhutdinov, and A. J. Smola · 2017
Later among the works it cites.
Babyai: A platform to study the sample efficiency of grounded language learning
M. Chevalier-Boisvert, D. Bahdanau, S. Lahlou, L. Willems, C. Saharia, T. H. Nguyen, and Y. Bengio · 2018
Later among the works it cites.
Minimalistic gridworld environment for openai gym
M. Chevalier-Boisvert, L. Willems, and S. Pal · 2018
Later among the works it cites.
Recurrent world models facilitate policy evolution
D. Ha and J. Schmidhuber · 2018
Later among the works it cites.
Distributed prioritized experience replay
D. Horgan, J. Quan, D. Budden, G. Barth-Maron, M. Hessel, H. van Hasselt, and D. Silver · 2018
Later among the works it cites.
Sparse attentive backtracking: Temporal credit assignment through reminding
N. R. Ke, A. Goyal, O. Bilaniuk, J. Binas, M. C. Mozer, C. Pal, and Y. Bengio · 2018
Later among the works it cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
Nervenet: Learning structured policy with graph neural networks
T. Wang, R. Liao, J. Ba, and S. Fidler · 2018
Later among the works it cites.
X. Wang, W. Xiong, H. Wang, and W. Y. Wang · 2018
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al · 2019
Later among the works it cites.
What is consciousness, and could machines have it?
S. Dehaene, H. Lau, and S. Kouider · 2020
Later among the works it cites.
Temporally abstract partial models
K. Khetarpal, Z. Ahmed, G. Comanici, and D. Precup · 2021
Closest in time.
Modeling event plausibility with consistent conceptual abstraction
I. Porada, K. Suleman, A. Trischler, and J. C. K. Cheung · 2021
Closest in time.