Fetching the paper…
Reading the bibliography…
Many important robotics problems are partially observable in the sense that a single visual or force-feedback measurement is insufficient to reconstruct the state.
Learning representations by back-propagating errors
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Learning policies for partially observable environments: Scaling up
M. L. Littman, A. R. Cassandra, and L. P. Kaelbling · 1995
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
L. P. Kaelbling, M. L. Littman, and A. R. Cassandra · 1998
Earlier work this paper cites.
Predictive representations of state
M. L. Littman and R. S. Sutton · 2002
Earlier work this paper cites.
Heuristic search value iteration for POMDPs
T. Smith and R. Simmons · 2003
Earlier work this paper cites.
Predictive state representations: A new theory for modeling dynamical systems
S. Singh, M. James, and M. Rudary · 2004
Earlier work this paper cites.
SARSOP: Efficient point-based POMDP planning by approximating optimally reachable belief spaces
H. Kurniawati, D. Hsu, and W. S. Lee · 2008
Earlier work this paper cites.
POMDPs for robotic tasks with mixed observability
S. C. Ong, S. W. Png, D. Hsu, and W. S. Lee · 2009
Earlier work this paper cites.
Monte-carlo planning in large POMDPs
D. Silver and J. Veness · 2010
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Lecture 6.5-RMSprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
DESPOT: Online POMDP planning with regularization
A. Somani, N. Ye, D. Hsu, and W. S. Lee · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
R. Pascanu, T. Mikolov, and Y. Bengio · 2013
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
K. Cho, B. Van Merriënboer, D. Bahdanau, and Y. Bengio · 2014
Earlier work this paper cites.
Deep recurrent q-learning for partially observable MDPs
M. Hausknecht and P. Stone · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
POMDP manipulation via trajectory optimization
N. A. Vien and M. Toussaint · 2015
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Learning from the memory of Atari 2600
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Later among the works it cites.
Learning to represent haptic feedback for partially-observable tasks
J. Sung, J. K. Salisbury, and A. Saxena · 2017
Later among the works it cites.
Asymmetric actor critic for image-based robot learning
L. Pinto, M. Andrychowicz, P. Welinder, W. Zaremba, and P. Abbeel · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
Learning to grasp under uncertainty using POMDPs
N. P. Garg, D. Hsu, and W. S. Lee · 2019
Later among the works it cites.
A floating-piston hydrostatic linear actuator and remote-direct-drive 2-dof gripper
E. Schwarm, K. M. Gravesmill, and J. P. Whitney · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Sygnowski and H. Michalewski · 2016
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
M. Jaderberg, V. Mnih, W. M. Czarnecki, T. Schaul, J. Z. Leibo, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Learning to navigate in complex environments
P. Mirowski, R. Pascanu, F. Viola, H. Soyer, A. J. Ballard, A. Banino, M. Denil, R. Goroshin, L. Sifre, K. Kavukcuoglu, et al · 2016
Cited alongside, same era.
Loss is its own reward: Self-supervision for reinforcement learning
E. Shelhamer, P. Mahmoudieh, M. Argus, and T. Darrell · 2016
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel · 2016
Cited alongside, same era.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
OpenAI baselines, 2017
P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, Y. Wu, and P. Zhokhov · 2017
Cited alongside, same era.
Reconciling λ \lambda -returns with experience replay
B. Daley and C. Amato · 2019
Later among the works it cites.
Learning to manipulate object collections using grounded state representations
M. Wilson and T. Hermans · 2020
Closest in time.
Learning by cheating
D. Chen, B. Zhou, V. Koltun, and P. Krähenbühl · 2020
Closest in time.
CURL: Contrastive unsupervised representations for reinforcement learning
A. Srinivas, M. Laskin, and P. Abbeel · 2020
Closest in time.
Learning complementary representations of the past using auxiliary tasks in partially observable reinforcement learning
A. Baisero and C. Amato · 2020
Closest in time.
Controlling contact-rich manipulation under partial observability
F. Wirnshofer, P. S. Schmitt, G. v. Wichert, and W. Burgard · 2020
Closest in time.
Learning force control for contact-rich manipulation tasks with rigid position-controlled robots
C. C. Beltran-Hernandez, D. Petit, I. G. Ramirez-Alpizar, T. Nishi, S. Kikuchi, T. Matsubara, and K. Harada · 2020
Closest in time.