Fetching the paper…
Reading the bibliography…
We introduce the value iteration network (VIN): a fully differentiable neural network with a `planning module' embedded within.
Dynamic Programming
R. Bellman · 1957
Earlier work this paper cites.
Neural network model for a mechanism of pattern recognition unaffected by shift in position- neocognitron
K. Fukushima · 1979
Earlier work this paper cites.
An on-line algorithm for dynamic reinforcement learning and planning in reactive environments
J. Schmidhuber · 1990
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Efficient learning in cellular simultaneous recurrent neural networks-the case of maze navigation problem
R. Ilin, R. Kozma, and P. J. Werbos · 2007
Earlier work this paper cites.
Apprenticeship learning using inverse reinforcement learning and gradient methods
G. Neu and C. Szepesvári · 2007
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
M. Deisenroth and C. E. Rasmussen · 2011
Earlier work this paper cites.
Hierarchical task and motion planning in the now
L. P. Kaelbling and T. Lozano-Pérez · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and A. Bagnell · 2011
Earlier work this paper cites.
Dynamic Programming and Optimal Control, Vol II
D. Bertsekas · 2012
Earlier work this paper cites.
Multi-column deep neural networks for image classification
D. Ciresan, U. Meier, and J. Schmidhuber · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. Hinton · 2012
Cited alongside, same era.
Lecture 6.5
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Cited alongside, same era.
Learning hierarchical features for scene labeling
C. Farabet, C. Couprie, L. Najman, and Y. LeCun · 2013
Cited alongside, same era.
Reinforcement learning with misspecified model classes
J. Joseph, A. Geramifard, J. W. Roberts, J. P. How, and N. Roy · 2013
Cited alongside, same era.
Embed to control: A locally linear latent dynamics model for control from raw images
M. Watter, J. Springenberg, J. Boedecker, and M. Riedmiller · 2015
Later among the works it cites.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio · 2015
Later among the works it cites.
Benchmarking deep reinforcement learning for continuous control
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel · 2016
Closest in time.
Guided policy search code implementation, 2016
C. Finn, M. Zhang, J. Fu, X. Tan, Z. McCarthy, E. Scharff, and S. Levine · 2016
Closest in time.
A machine learning approach to visual perception of forest trails for mobile robots
A. Giusti et al · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep learning for real-time atari game play using offline monte-carlo tree search planning
X. Guo, S. Singh, H. Lee, R. L. Lewis, and X. Wang · 2014
Cited alongside, same era.
Learning neural network policies with guided policy search under unknown dynamics
S. Levine and P. Abbeel · 2014
Cited alongside, same era.
Fully convolutional networks for semantic segmentation
J. Long, E. Shelhamer, and T. Darrell · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Rusu, J. Veness, M. Bellemare, A. Graves, M. Riedmiller, A. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Cited alongside, same era.
X. Guo, S. Singh, R. Lewis, and H. Lee · 2016
Closest in time.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Closest in time.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Closest in time.
Webnav: A new large-scale task for natural language based sequential decision making
R. Nogueira and K. Cho · 2016
Closest in time.
Theano: A Python framework for fast computation of mathematical expressions
Theano Development Team · 2016
Closest in time.