Fetching the paper…
Reading the bibliography…
Policy learning for partially observed control tasks requires policies that can remember salient information from past observations.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,”
1997
Earlier work this paper cites.
N. Meuleau, L. Peshkin, K.-E. Kim, and L. P. Kaelbling, “Learning finite-state controllers for partially observable environments,” in
1999
Earlier work this paper cites.
L. Peshkin, N. Meuleau, and L. Kaelbling, “Learning policies with external memory,”
2001
Earlier work this paper cites.
J. A. Bagnell and J. Schneider, “Covariant policy search,” in
2003
Earlier work this paper cites.
D. Wierstra, A. Foerster, J. Peters, and J. Schmidhuber, “Solving deep memory POMDPs with recurrent policy gradients,” in
2007
Earlier work this paper cites.
J. Peters and S. Schaal, “Applying the episodic natural actor-critic architecture to motor primitive learning,” in
2007
Earlier work this paper cites.
M. Toussaint, L. Charlin, and P. Poupart, “Hierarchical POMDP controller optimization by likelihood maximization.” in
2008
Earlier work this paper cites.
S. Ross, J. Pineau, S. Paquet, and B. Chaib-Draa, “Online planning algorithms for POMDPs,”
2008
Earlier work this paper cites.
E. Brunskill, L. P. Kaelbling, T. Lozano-Perez, and N. Roy, “Continuous-state POMDPs with hybrid dynamics.” in
2008
Cited alongside, same era.
J. Peters and S. Schaal, “Reinforcement learning of motor skills with policy gradients,”
2008
Cited alongside, same era.
J. Kober and J. Peters, “Learning motor primitives for robotics,” in
2009
Cited alongside, same era.
J. Peters, K. Mülling, and Y. Altün, “Relative entropy policy search,” in
2010
Cited alongside, same era.
M. Deisenroth and C. E. Rasmussen, “PILCO: A model-based and data-efficient approach to policy search,” in
2011
Cited alongside, same era.
S. Ross, G. Gordon, and A. Bagnell, “A reduction of imitation learning and structured prediction to no-regret online learning,”
M. Deisenroth, G. Neumann, and J. Peters, “A survey on policy search for robotics,”
2013
Later among the works it cites.
S. Levine and P. Abbeel, “Learning neural network policies with guided policy search under unknown dynamics,” in
2014
Later among the works it cites.
2014
Later among the works it cites.
S. Levine and V. Koltun, “Learning complex neural network policies with trajectory optimization,” in
2014
Later among the works it cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in
2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2011
Cited alongside, same era.
R. Pascanu, T. Mikolov, and Y. Bengio, “On the difficulty of training recurrent neural networks,”
2012
Cited alongside, same era.
G. Shani, J. Pineau, and R. Kaplow, “A survey of point-based POMDP solvers,”
2013
Cited alongside, same era.
M. Watter, J. T. Springenberg, J. Boedecker, and M. Riedmiller, “Embed to control: A locally linear latent dynamics model for control from raw images,” in
2015
Closest in time.
S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,”
2015
Closest in time.
S. Levine, N. Wagener, and P. Abbeel, “Learning contact-rich manipulation skills with guided policy search,” 2015
2015
Closest in time.