Fetching the paper…
Reading the bibliography…
Addressing uncertainty is critical for autonomous systems to robustly adapt to the real world.
The complexity of Markov decision processes
Christos Papadimitriou and John Tsitsiklis · 1987
Earlier work this paper cites.
Learning policies for partially observable environments: Scaling up
Michael Littman, Anthony Cassandra, and Leslie Pack Kaelbling · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael Littman, and Anthony Cassandra · 1998
Earlier work this paper cites.
Variable resolution discretization for high-accuracy solutions of optimal control problems
Remi Munos and Andrew Moore · 1999
Earlier work this paper cites.
A Bayesian framework for reinforcement learning
Malcolm Strens · 2000
Earlier work this paper cites.
Robust reinforcement learning
Jun Morimoto and Kenji Doya · 2001
Earlier work this paper cites.
The optimal reward baseline for gradient-based reinforcement learning
Lex Weaver and Nigel Tao · 2001
Earlier work this paper cites.
Scaling internal-state policy-gradient methods for POMDPs
Douglas Aberdeen and Jonathan Baxter · 2002
Earlier work this paper cites.
Optimal Learning: Computational procedures for Bayes-adaptive Markov decision processes
Michael O’Gordon Duff and Andrew Barto · 2002
Earlier work this paper cites.
A (revised) survey of approximate methods for solving partially observable markov decision processes
Douglas Aberdeen · 2003
Earlier work this paper cites.
Point-based value iteration: An anytime algorithm for POMDPs
Joelle Pineau, Geoff Gordon, Sebastian Thrun, et al · 2003
Earlier work this paper cites.
Perseus: Randomized point-based value iteration for POMDPs
Matthijs TJ Spaan and Nikos Vlassis · 2005
Earlier work this paper cites.
An analytic solution to discrete bayesian reinforcement learning
Pascal Poupart, Nikos Vlassis, Jesse Hoey, and Kevin Regan · 2006
Earlier work this paper cites.
H-infinity optimal control and related minimax design problems: a dynamic game approach
Tamer Başar and Pierre Bernhard · 2008
Earlier work this paper cites.
SARSOP: Efficient point-based POMDP planning by approximating optimally reachable belief spaces
Hanna Kurniawati, David Hsu, and Wee Sun Lee · 2008
Cited alongside, same era.
Bayes-adaptive POMDPs
Stephane Ross, Brahim Chaib-draa, and Joelle Pineau · 2008
Cited alongside, same era.
Near-Bayesian exploration in polynomial time
Zico Kolter and Andrew Ng · 2009
Cited alongside, same era.
Planning under uncertainty for robotic tasks with mixed observability
Sylvie CW Ong, Shao Wei Png, David Hsu, and Wee Sun Lee · 2010
Cited alongside, same era.
Belief space planning assuming maximum likelihood observations
Robert Platt, Russ Tedrake, Leslie Pack Kaelbling, and Tomas Lozano-Perez · 2010
Cited alongside, same era.
Monte-carlo planning in large POMDPs
David Silver and Joel Veness · 2010
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Later among the works it cites.
Preparing for the unknown: Learning a universal policy with online system identification
Wenhao Yu, Jie Tan, C. Karen Liu, and Greg Turk · 2015
Later among the works it cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Later among the works it cites.
POMDP-lite for robust robot planning under uncertainty
Min Chen, Emilio Frazzoli, David Hsu, and Wee Sun Lee · 2016
Later among the works it cites.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Later among the works it cites.
QMDP-Net: Deep learning for planning under partial observability
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Efficient Bayes-adaptive reinforcement learning using sample-based search
Arthur Guez, David Silver, and Peter Dayan · 2012
Cited alongside, same era.
MuJoCo: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
Motion planning under uncertainty using iterative local optimization in belief space
Jur van den Berg, Sachin Patil, and Ron Alterovitz · 2012
Cited alongside, same era.
Monte Carlo Bayesian reinforcement learning
Yi Wang, Kok Sung Won, David Hsu, and Wee Sun Lee · 2012
Cited alongside, same era.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Cited alongside, same era.
A survey of point-based POMDP solvers
Guy Shani, Joelle Pineau, and Robert Kaplow · 2013
Cited alongside, same era.
Peter Karkus, David Hsu, and Wee Sun Lee · 2017
Later among the works it cites.
Robust adversarial reinforcement learning
Lerrel Pinto, James Davidson, Rahul Sukthankar, and Abhinav Gupta · 2017
Later among the works it cites.
EPOpt: Learning robust neural network policies using model ensembles
Aravind Rajeswaran, Sarvjeet Ghotra, Balaraman Ravindran, and Sergey Levine · 2017
Later among the works it cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Emergent complexity via multi-agent competition
Trapit Bansal, Jakub Pachocki, Szymon Sidor, Ilya Sutskever, and Igor Mordatch · 2018
Closest in time.
Deep variational reinforcement learning for pomdps
Maximilian Igl, Luisa Zintgraf, Tuan Anh Le, Frank Wood, and Shimon Whiteson · 2018
Closest in time.
Robust deep reinforcement learning with adversarial attacks
Anay Pattanaik, Zhenyi Tang, Shuijing Liu, Gautham Bommannan, and Girish Chowdhary · 2018
Closest in time.
Sim-to-real transfer of robotic control with dynamics randomization
Xue Bin Peng, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel · 2018
Closest in time.
Online algorithms for POMDPs with continuous state, action, and observation spaces
Zachary Sunberg and Mykel Kochenderfer · 2018
Closest in time.