Fetching the paper…
Reading the bibliography…
Reinforcement learning in partially observable domains is challenging due to the lack of observable state information.
Optimal control of markov decision processes with incomplete state estimation
K. J. Astrom · 1965
Earlier work this paper cites.
An image synthesizer
K. Perlin · 1985
Earlier work this paper cites.
Learning without state-estimation in partially observable markovian decision processes
S. P. Singh, T. Jaakkola, and M. I. Jordan · 1994
Earlier work this paper cites.
A framework for behavioural cloning
M. Bain and C. Sammut · 1995
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
L. P. Kaelbling, M. L. Littman, and A. R. Cassandra · 1998
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Learning monocular reactive uav control in cluttered natural environments
S. Ross, N. Melik-Barkhudarov, K. S. Shankar, A. Wendel, D. Dey, J. A. Bagnell, and M. Hebert · 2013
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
K. Cho, B. Van Merriënboer, D. Bahdanau, and Y. Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
A machine learning approach to visual perception of forest trails for mobile robots
A. Giusti, J. Guzzi, D. C. Cireşan, F.-L. He, J. P. Rodríguez, F. Fontana, M. Faessler, C. Forster, J. Schmidhuber, G. Di Caro, et al · 2015
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
End to end learning for self-driving cars
M. Bojarski, D. Del Testa, D. Dworakowski, B. Firner, B. Flepp, P. Goyal, L. D. Jackel, M. Monfort, U. Muller, J. Zhang, et al · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
Pybullet, a python module for physics simulation for games, robotics and machine learning
E. Coumans and Y. Bai · 2016
Cited alongside, same era.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
M. Vecerik, T. Hester, J. Scholz, F. Wang, O. Pietquin, B. Piot, N. Heess, T. Rothörl, T. Lampe, and M. Riedmiller · 2017
Cited alongside, same era.
Transferring end-to-end visuomotor control from simulation to real world for a multi-stage task
S. James, A. J. Davison, and E. Johns · 2017
Soft actor-critic for discrete action settings
P. Christodoulou · 2019
Later among the works it cites.
Learning belief representations for imitation learning in pomdps
T. Gangwani, J. Lehman, Q. Liu, and J. Peng · 2020
Later among the works it cites.
Memory-based deep reinforcement learning for pomdps
L. Meng, R. Gorbet, and D. Kulić · 2021
Later among the works it cites.
Belief-grounded networks for accelerated robot learning under partial observability
H. Nguyen, B. Daley, X. Song, C. Amato, and R. Platt · 2021
Later among the works it cites.
Robust asymmetric learning in pomdps
A. Warrington, J. W. Lavington, A. Scibior, M. Schmidt, and F. Wood · 2021
Later among the works it cites.
Bridging the imitation gap by adaptive insubordination
L. Weihs, U. Jain, I.-J. Liu, J. Salvador, S. Lazebnik, A. Kembhavi, and A. Schwing · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep variational reinforcement learning for pomdps
M. Igl, L. Zintgraf, T. A. Le, F. Wood, and S. Whiteson · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Asymmetric actor critic for image-based robot learning
L. Pinto, M. Andrychowicz, P. Welinder, W. Zaremba, and P. Abbeel · 2018
Cited alongside, same era.
Deep q-learning from demonstrations
T. Hester, M. Vecerik, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, D. Horgan, J. Quan, A. Sendonaris, I. Osband, et al · 2018
Cited alongside, same era.
Soft actor-critic algorithms and applications
T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V. Kumar, H. Zhu, A. Gupta, P. Abbeel, et al · 2018
Cited alongside, same era.
Minimalistic gridworld environment for openai gym
M. Chevalier-Boisvert, L. Willems, and S. Pal · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. Hoof, and D. Meger · 2018
Cited alongside, same era.
Later among the works it cites.
Pomdp robot domains
H. Nguyen · 2021
Later among the works it cites.
Demonstration actor critic
G. Liu, L. Zhao, P. Zhang, J. Bian, T. Qin, N. Yu, and T.-Y. Liu · 2021
Later among the works it cites.
Stable-baselines3: Reliable reinforcement learning implementations
A. Raffin, A. Hill, A. Gleave, A. Kanervisto, M. Ernestus, and N. Dormann · 2021
Later among the works it cites.
Partially observable markov decision processes and robotics
H. Kurniawati · 2022
Closest in time.
Hierarchical reinforcement learning under mixed observability
H. Nguyen, Z. Yang, A. Baisero, X. Ma, R. Platt, and C. Amato · 2022
Closest in time.
Asymmetric dqn for partially observable reinforcement learning
A. Baisero, B. Daley, and C. Amato · 2022
Closest in time.
Unbiased asymmetric reinforcement learning under partial observability
A. Baisero and C. Amato · 2022
Closest in time.
Recurrent model-free rl can be a strong baseline for many pomdps
T. Ni, B. Eysenbach, and R. Salakhutdinov · 2022
Closest in time.
Bulletarm: An open-source robotic manipulation benchmark and learning framework
D. Wang, C. Kohler, X. Zhu, M. Jia, and R. Platt · 2022
Closest in time.