Fetching the paper…
Reading the bibliography…
This paper introduces the QMDP-net, a neural network architecture for planning under partial observability.
The complexity of Markov decision processes
C. H. Papadimitriou and J. N. Tsitsiklis · 1987
Earlier work this paper cites.
Learning policies for partially observable environments: Scaling up
M. L. Littman, A. R. Cassandra, and L. P. Kaelbling · 1995
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
A. Y. Ng, D. Harada, and S. Russell · 1999
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
J. Baxter and P. L. Bartlett · 2001
Earlier work this paper cites.
Predictive representations of state
M. L. Littman, R. S. Sutton, and S. Singh · 2002
Earlier work this paper cites.
Policy search by dynamic programming
J. A. Bagnell, S. Kakade, A. Y. Ng, and J. G. Schneider · 2003
Earlier work this paper cites.
A robot that reinforcement-learns to identify and memorize important previous observations
B. Bakker, V. Zhumatiy, G. Gruener, and J. Schmidhuber · 2003
Earlier work this paper cites.
The robotics data set repository (radish), 2003
A. Howard and N. Roy · 2003
Earlier work this paper cites.
Applying metric-trees to belief-point POMDPs
J. Pineau, G. J. Gordon, and S. Thrun · 2003
Earlier work this paper cites.
Model-based online learning of POMDPs
G. Shani, R. I. Brafman, and S. E. Shimony · 2005
Earlier work this paper cites.
Perseus: Randomized point-based value iteration for POMDPs
M. T. Spaan and N. Vlassis · 2005
Earlier work this paper cites.
Grasping POMDPs
K. Hsiao, L. P. Kaelbling, and T. Lozano-Pérez · 2007
Cited alongside, same era.
Sarsop: Efficient point-based POMDP planning by approximating optimally reachable belief spaces
H. Kurniawati, D. Hsu, and W. S. Lee · 2008
Cited alongside, same era.
Monte carlo value iteration for continuous-state POMDPs
H. Bai, D. Hsu, W. S. Lee, and V. A. Ngo · 2010
Cited alongside, same era.
Monte-carlo planning in large POMDPs
D. Silver and J. Veness · 2010
Cited alongside, same era.
Closing the learning-planning loop with predictive state representations
B. Boots, S. M. Siddiqi, and G. J. Gordon · 2011
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
Deep recurrent Q-learning for partially observable MDPs
M. J. Hausknecht and P. Stone · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Later among the works it cites.
Convolutional LSTM network: A machine learning approach for precipitation nowcasting
S. Xingjian, Z. Chen, H. Wang, D.-Y. Yeung, W.-k. Wong, and W.-c. Woo · 2015
Later among the works it cites.
Backprop kf: Learning discriminative deterministic state estimators
T. Haarnoja, A. Ajay, S. Levine, and P. Abbeel · 2016
Later among the works it cites.
End-to-end learnable histogram filters
R. Jonschkowski and O. Brock · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lecture 6.5 - rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
3D convolutional neural networks for human action recognition
S. Ji, W. Xu, M. Yang, and K. Yu · 2013
Cited alongside, same era.
A survey of point-based POMDP solvers
G. Shani, J. Pineau, and R. Kaplow · 2013
Cited alongside, same era.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, et al · 2015
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al
Cited in the paper.
P. Mirowski, R. Pascanu, F. Viola, H. Soyer, A. Ballard, A. Banino, M. Denil, R. Goroshin, L. Sifre, K. Kavukcuoglu, et al · 2016
Later among the works it cites.
Reinforcement learning via recurrent convolutional neural networks
T. Shankar, S. K. Dwivedy, and P. Guha · 2016
Later among the works it cites.
Value iteration networks
A. Tamar, S. Levine, P. Abbeel, Y. Wu, and G. Thomas · 2016
Later among the works it cites.
Cognitive mapping and planning for visual navigation
S. Gupta, J. Davidson, S. Levine, R. Sukthankar, and J. Malik · 2017
Closest in time.
Path integral networks: End-to-end differentiable optimal control
M. Okada, L. Rigazio, and T. Aoshima · 2017
Closest in time.
Despot: Online POMDP planning with regularization
N. Ye, A. Somani, D. Hsu, and W. S. Lee · 2017
Closest in time.