Fetching the paper…
Reading the bibliography…
The partially observable Markov decision process (POMDP) provides a principled general model for planning under uncertainty.
E. Sondik, “The optimal control of partially observable Markov processes,” Ph.D. dissertation, Stanford University, Stanford, California, USA, 1971
1971
Earlier work this paper cites.
C. H. Papadimitriou and J. N. Tsitsiklis, “The complexity of markov decision processes,” Mathematics of operations research , vol. 12, no. 3, pp. 441–450, 1987
1987
Earlier work this paper cites.
L. Kaelbling, M. Littman, and A. Cassandra, “Planning and acting in partially observable stochastic domains,” Artificial Intelligence , vol. 101, no. 1–2, pp. 99–134, 1998
1998
Earlier work this paper cites.
N. Roy and S. Thrun, “Coastal navigation with mobile robots,” in Advances in Neural Information Processing Systems . The MIT Press, 1999, vol. 12, pp. 1043–1049
1999
Earlier work this paper cites.
M. Kearns and S. Singh, “Near-optimal reinforcement learning in polynomial time,” Machine Learning , vol. 49, no. 2-3, pp. 209–232, 2002
2002
Earlier work this paper cites.
T. Smith and R. Simmons, “Point-based POMDP algorithms: Improved analysis and implementation,” in Proc. Conf. on Uncertainty in Artificial Intelligence , 2005
2005
Earlier work this paper cites.
S. Thrun, W. Burgard, and D. Fox, Probabilistic Robotics . The MIT Press, 2005
2005
Earlier work this paper cites.
L. Kocsis and C. Szepesvári, “Bandit based monte-carlo planning,” in Machine Learning: ECML 2006 . Springer, 2006, pp. 282–293
2006
Earlier work this paper cites.
P. Poupart, N. Vlassis, J. Hoey, and K. Regan, “An analytic solution to discrete bayesian reinforcement learning,” in Proceedings of the 23rd international conference on Machine learning . ACM, 2006
2006
Earlier work this paper cites.
A. L. Strehl, L. Li, and M. L. Littman, “Incremental model-based learners with formal learning-time guarantees,” in UAI, Proceedings of the 22nd Conference in Uncertainty in Artificial Intelligence , 2006
2006
Earlier work this paper cites.
A. Fern, S. Natarajan, K. Judah, and P. Tadepalli, “A decision-theoretic model of assistance,” in Proc. AAAI Conf. on Artificial Intelligence , 2007
2007
Cited alongside, same era.
K. Hsiao, L. Kaelbling, and T. Lozano-Pérez, “Grasping POMDPs,” in Proc. IEEE Int. Conf. on Robotics & Automation , 2007
2007
Cited alongside, same era.
H. Kurniawati, D. Hsu, and W. Lee, “SARSOP: Efficient point-based POMDP planning by approximating optimally reachable belief spaces,” in Proc. Robotics: Science & Systems , 2008
2008
Cited alongside, same era.
S. Ross, J. Pineau, S. Paquet, and B. Chaib-Draa, “Online planning algorithms for POMDPs,” J. Artificial Intelligence Research , vol. 32, no. 1, pp. 663–704, 2008
2008
Cited alongside, same era.
J. Asmuth, L. Li, M. L. Littman, A. Nouri, and D. Wingate, “A bayesian sampling approach to exploration in reinforcement learning,” in Proceedings of the Twenty-Fifth Conference on Uncertainty in Artificial Intelligence , 2009
T. Bandyopadhyay, K. Won, E. Frazzoli, D. Hsu, W. Lee, and D. Rus, “Intention-aware motion planning,” in Algorithmic Foundations of Robotics X—Proc. Int. Workshop on the Algorithmic Foundations of Robotics (WAFR) , 2012
2012
Later among the works it cites.
H. Bai, D. Hsu, and W. Lee, “Planning how to learn,” in Proc. IEEE Int. Conf. on Robotics & Automation , 2013
2013
Later among the works it cites.
2013
Later among the works it cites.
E. Rohmer, S. P. Singh, and M. Freese, “V-rep: A versatile and scalable robot simulation framework,” in Intelligent Robots and Systems (IROS), 2013 IEEE/RSJ International Conference on . IEEE, 2013, pp. 1321–1326
2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2009
Cited alongside, same era.
J. Z. Kolter and A. Y. Ng, “Near-bayesian exploration in polynomial time,” in Proceedings of the 26th Annual International Conference on Machine Learning , 2009
2009
Cited alongside, same era.
S. C. Ong, S. W. Png, D. Hsu, and W. S. Lee, “Planning under uncertainty for robotic tasks with mixed observability,” The International Journal of Robotics Research , vol. 29, no. 8, pp. 1053–1068, 2010
2010
Cited alongside, same era.
D. Silver and J. Veness, “Monte-Carlo planning in large POMDPs,” in Advances in Neural Information Processing Systems , 2010
2010
Cited alongside, same era.
J. Sorg, S. P. Singh, and R. L. Lewis, “Variance-based rewards for approximate bayesian reinforcement learning,” in UAI, Proceedings of the Twenty-Sixth Conference on Uncertainty in Artificial Intelligence , 2010
2010
Cited alongside, same era.
A. Somani, N. Ye, D. Hsu, and W. Lee, “DESPOT: Online POMDP planning with regularization,” in Advances in Neural Information Processing Systems , 2013
2013
Later among the works it cites.
H. Bai, S. Cai, D. Hsu, and W. Lee, “Intention-aware online POMDP planning for autonomous driving in a crowd,” in Proc. IEEE Int. Conf. on Robotics & Automation , 2015
2015
Later among the works it cites.
M. Koval, N. Pollard, and S. Srinivasa, “Pre- and post-contact policy decomposition for planar contact manipulation under uncertainty,” Int. J. Robotics Research , 2015
2015
Later among the works it cites.
S. Nikolaidis, R. Ramakrishnan, K. Gu, and J. Shah, “Efficient model learning from joint-action demonstrations for human-robot collaborative tasks,” in Proc. ACM/IEEE Int. Conf. on Human-Robot Interaction , 2015
2015
Later among the works it cites.
K. Seiler, H. Kurniawati, and S. Singh, “An online and approximate solver for pomdps with continuous action space,” in Proc. IEEE Int. Conf. on Robotics & Automation , 2015
2015
Later among the works it cites.