Fetching the paper…
Reading the bibliography…
We present Value Propagation (VProp), a set of parameter-efficient differentiable planning modules built on Value Iteration which can successfully be trained using reinforcement learning to solve unseen tasks, has the capability to generalize to larger map sizes, and can learn to navigate in dynamic environments.
Q-learning
C. J. Watkins and P. Dayan · 1992
Earlier work this paper cites.
A modern approach
S. Russell, P. Norvig, and A. Intelligence · 1995
Earlier work this paper cites.
Reinforcement learning: A survey
L. P. Kaelbling, M. L. Littman, and A. W. Moore · 1996
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, Y. Mansour, et al · 1999
Earlier work this paper cites.
Actor-critic algorithms
V. R. Konda and J. N. Tsitsiklis · 2000
Earlier work this paper cites.
Dynamic Programming and Optimal Control, Vol. II
D. P. Bertsekas · 2012
Earlier work this paper cites.
Counterfactual reasoning about intent for interactive navigation in dynamic environments
A. Bordallo, F. Previtali, N. Nardelli, and S. Ramamoorthy · 2015
Earlier work this paper cites.
Mazebase: A sandbox for learning from games
S. Sukhbaatar, A. Szlam, G. Synnaeve, S. Chintala, and R. Fergus · 2015
Earlier work this paper cites.
Playing doom with slam-augmented deep reinforcement learning
S. Bhatti, A. Desmaison, O. Miksik, N. Nardelli, N. Siddharth, and P. H. Torr · 2016
Earlier work this paper cites.
Learning to navigate in complex environments
P. Mirowski, R. Pascanu, F. Viola, H. Soyer, A. Ballard, A. Banino, M. Denil, R. Goroshin, L. Sifre, K. Kavukcuoglu, et al · 2016
Earlier work this paper cites.
The predictron: End-to-end learning and planning
D. Silver, H. van Hasselt, M. Hessel, T. Schaul, A. Guez, T. Harley, G. Dulac-Arnold, D. Reichert, N. Rabinowitz, A. Barreto, et al · 2016
Cited alongside, same era.
Torchcraft: a library for machine learning research on real-time strategy games
G. Synnaeve, N. Nardelli, A. Auvolat, S. Chintala, T. Lacroix, Z. Lin, F. Richoux, and N. Usunier · 2016
Cited alongside, same era.
Value iteration networks
A. Tamar, S. Levine, P. Abbeel, Y. WU, and G. Thomas · 2016
Cited alongside, same era.
N. Usunier, G. Synnaeve, Z. Lin, and S. Chintala · 2016
Cited alongside, same era.
Desire: Distant future prediction in dynamic scenes with interacting agents
N. Lee, W. Choi, P. Vernaza, C. B. Choy, P. H. Torr, and M. Chandraker · 2017
Later among the works it cites.
Generalized value iteration networks: Life beyond lattices
S. Niu, S. Chen, H. Guo, C. Targonski, M. C. Smith, and J. Kovačević · 2017
Later among the works it cites.
J. Oh, S. Singh, and H. Lee · 2017
Later among the works it cites.
Automatic differentiation in pytorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer · 2017
Later among the works it cites.
Cooperative motion planning for non-holonomic agents with value iteration networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Wang, V. Bapst, N. Heess, V. Mnih, R. Munos, K. Kavukcuoglu, and N. de Freitas · 2016
Cited alongside, same era.
Unifying task specification in reinforcement learning
M. White · 2016
Cited alongside, same era.
Treeqn and atreec: Differentiable tree planning for deep reinforcement learning
G. Farquhar, T. Rocktäschel, M. Igl, and S. Whiteson · 2017
Cited alongside, same era.
Stabilising experience replay for deep multi-agent reinforcement learning
J. Foerster, N. Nardelli, G. Farquhar, P. Torr, P. Kohli, S. Whiteson, et al · 2017
Cited alongside, same era.
Learning generalized reactive policies using deep neural networks
E. Groshev, A. Tamar, S. Srivastava, and P. Abbeel · 2017
Cited alongside, same era.
Cognitive mapping and planning for visual navigation
S. Gupta, J. Davidson, S. Levine, R. Sukthankar, and J. Malik · 2017
Cited alongside, same era.
Memory augmented control networks
A. Khan, C. Zhang, N. Atanasov, K. Karydis, V. Kumar, and D. D. Lee · 2017
Cited alongside, same era.
E. Rehder, M. Naumann, N. O. Salscheider, and C. Stiller · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
T. Weber, S. Racanière, D. P. Reichert, L. Buesing, A. Guez, D. J. Rezende, A. P. Badia, O. Vinyals, N. Heess, Y. Li, et al · 2017
Later among the works it cites.
J. Zhang, L. Tai, J. Boedecker, W. Burgard, and M. Liu · 2017
Later among the works it cites.
Vector-based navigation using grid-like representations in artificial agents
A. Banino, C. Barry, B. Uria, C. Blundell, T. Lillicrap, P. Mirowski, A. Pritzel, M. J. Chadwick, T. Degris, J. Modayil, et al · 2018
Closest in time.
The bottleneck simulator: A model-based deep reinforcement learning approach
I. V. Serban, C. Sankar, M. Pieper, J. Pineau, and Y. Bengio · 2018
Closest in time.
Universal planning networks: Learning generalizable representations for visuomotor control
A. Srinivas, A. Jabri, P. Abbeel, S. Levine, and C. Finn · 2018
Closest in time.