Fetching the paper…
Reading the bibliography…
In recent years, deep generative models have been shown to 'imagine' convincing high-dimensional observations such as images, audio, and even video, learning directly from raw data.
Learning and executing generalized robot plans
R. E. Fikes, P. E. Hart, and N. J. Nilsson · 1972
Earlier work this paper cites.
Shakey the robot
N. J. Nilsson · 1984
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
The information bottleneck method
N. Tishby, F. C. Pereira, and W. Bialek · 2000
Earlier work this paper cites.
Dynamic abstraction in reinforcement learning via clustering
S. Mannor, I. Menache, A. Hoze, and U. Klein · 2004
Earlier work this paper cites.
Dynamic programming and optimal control: Volume II
D. Bertsekas · 2005
Earlier work this paper cites.
Dynamic catalog mailing policies
D. I. Simester, P. Sun, and J. N. Tsitsiklis · 2006
Earlier work this paper cites.
Proto-value functions: A laplacian framework for learning representation and control in markov decision processes
S. Mahadevan and M. Maggioni · 2007
Earlier work this paper cites.
Artificial Intelligence - A Modern Approach (3. internat. ed.)
S. J. Russell and P. Norvig · 2010
Earlier work this paper cites.
Next-generation airborne collision avoidance system
M. J. Kochenderfer, J. E. Holland, and J. P. Chryssanthacopoulos · 2012
Earlier work this paper cites.
Simulation as an engine of physical scene understanding
P. W. Battaglia, J. B. Hamrick, and J. B. Tenenbaum · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Past-future information bottleneck for linear feedback systems
N. Amir, S. Tiomkin, and N. Tishby · 2015
Earlier work this paper cites.
A recurrent latent variable model for sequential data
J. Chung, K. Kastner, L. Dinh, K. Goel, A. C. Courville, and Y. Bengio · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Rusu, J. Veness, M. Bellemare, A. Graves, M. Riedmiller, A. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
J. Oh, X. Guo, H. Lee, R. L. Lewis, and S. Singh · 2015
Earlier work this paper cites.
Unsupervised representation learning with deep convolutional generative adversarial networks
A. Radford, L. Metz, and S. Chintala · 2015
Earlier work this paper cites.
Tractability of planning with loops
S. Srivastava, S. Zilberstein, A. Gupta, P. Abbeel, and S. Russell · 2015
Cited alongside, same era.
The 2014 international planning competition: Progress and trends
M. Vallati, L. Chrpa, M. Grześ, T. L. McCluskey, M. Roberts, S. Sanner, et al · 2015
Cited alongside, same era.
Embed to control: A locally linear latent dynamics model for control from raw images
M. Watter, J. Springenberg, J. Boedecker, and M. Riedmiller · 2015
Cited alongside, same era.
Learning to poke by poking: Experiential learning of intuitive physics
P. Agrawal, A. V. Nair, P. Abbeel, J. Malik, and S. Levine · 2016
Cited alongside, same era.
Spatio-temporal abstractions in reinforcement learning through neural encoding
N. Baram, T. Zahavy, and S. Mannor · 2016
Cited alongside, same era.
InfoGAN: Interpretable representation learning by information maximizing generative adversarial nets
Model-agnostic meta-learning for fast adaptation of deep networks
C. Finn, P. Abbeel, and S. Levine · 2017
Later among the works it cites.
Deep visual foresight for planning robot motion
C. Finn and S. Levine · 2017
Later among the works it cites.
Schema networks: Zero-shot transfer with a generative causal model of intuitive physics
K. Kansky, T. Silver, D. A. Mély, M. Eldawy, M. Lázaro-Gredilla, X. Lou, N. Dorfman, S. Sidor, S. Phoenix, and D. George · 2017
Later among the works it cites.
Progressive growing of gans for improved quality, stability, and variation
T. Karras, T. Aila, S. Laine, and J. Lehtinen · 2017
Later among the works it cites.
The eigenoption-critic framework
M. Liu, M. C. Machado, G. Tesauro, and M. Campbell · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Chen, Y. Duan, R. Houthooft, J. Schulman, I. Sutskever, and P. Abbeel · 2016
Cited alongside, same era.
J. Donahue, P. Krähenbühl, and T. Darrell · 2016
Cited alongside, same era.
RL 2 \mbox{RL}^{2} : Fast reinforcement learning via slow reinforcement learning
Y. Duan, J. Schulman, X. Chen, P. L. Bartlett, I. Sutskever, and P. Abbeel · 2016
Cited alongside, same era.
Guided cost learning: Deep inverse optimal control via policy optimization
C. Finn, S. Levine, and P. Abbeel · 2016
Cited alongside, same era.
Deep spatial autoencoders for visuomotor learning
C. Finn, X. Y. Tan, Y. Duan, T. Darrell, S. Levine, and P. Abbeel · 2016
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
E. Jang, S. Gu, and B. Poole · 2016
Cited alongside, same era.
Composing graphical models with neural networks for structured representations and fast inference
M. Johnson, D. K. Duvenaud, A. Wiltschko, R. P. Adams, and S. R. Datta · 2016
Cited alongside, same era.
M. C. Machado, M. G. Bellemare, and M. Bowling · 2017
Later among the works it cites.
Combining self-supervised learning and imitation for vision-based rope manipulation
A. Nair, D. Chen, P. Agrawal, P. Isola, P. Abbeel, J. Malik, and S. Levine · 2017
Later among the works it cites.
Value prediction network
J. Oh, S. Singh, and H. Lee · 2017
Later among the works it cites.
A simple neural network module for relational reasoning
A. Santoro, D. Raposo, D. G. Barrett, M. Malinowski, R. Pascanu, P. Battaglia, and T. Lillicrap · 2017
Later among the works it cites.
Safer classification by synthesis
W. Wang, A. Wang, A. Tamar, X. Chen, and P. Abbeel · 2017
Later among the works it cites.
N. Watters, A. Tacchetti, T. Weber, R. Pascanu, P. Battaglia, and D. Zoran · 2017
Later among the works it cites.
Learning to see physics via visual de-animation
J. Wu, E. Lu, P. Kohli, B. Freeman, and J. Tenenbaum · 2017
Later among the works it cites.
Toward multimodal image-to-image translation
J.-Y. Zhu, R. Zhang, D. Pathak, T. Darrell, A. A. Efros, O. Wang, and E. Shechtman · 2017
Later among the works it cites.
Efficient model-based deep reinforcement learning with variational state tabulation
D. Corneil, W. Gerstner, and J. Brea · 2018
Closest in time.
From skills to symbols: Learning symbolic representations for abstract high-level planning
G. Konidaris, L. P. Kaelbling, and T. Lozano-Perez · 2018
Closest in time.
Temporal difference models: Model-free deep RL for model-based control
V. Pong, S. Gu, M. Dalal, and S. Levine · 2018
Closest in time.
Learning by playing-solving sparse reward tasks from scratch
M. Riedmiller, R. Hafner, T. Lampe, M. Neunert, J. Degrave, T. Van de Wiele, V. Mnih, N. Heess, and J. T. Springenberg · 2018
Closest in time.
Semi-parametric topological memory for navigation
N. Savinov, A. Dosovitskiy, and V. Koltun · 2018
Closest in time.
A. Srinivas, A. Jabri, P. Abbeel, S. Levine, and C. Finn · 2018
Closest in time.