Fetching the paper…
Reading the bibliography…
Learning to solve long horizon temporally extended tasks with reinforcement learning has been a challenge for several years now.
Strips: A new approach to the application of theorem proving to problem solving
Fikes, R. E.; and Nilsson, N. J. 1971 · 1971
Earlier work this paper cites.
PDDL— The Planning Domain Definition Language
Aeronautiques, C.; Howe, A.; Knoblock, C.; McDermott, I. D.; Ram, A.; Veloso, M.; Weld, D.; SRI, D. W.; Barrett, A.; Christianson, D.; et al. 1998 · 1998
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Sutton, R. S.; Precup, D.; and Singh, S. 1999 · 1999
Earlier work this paper cites.
Grounding language to autonomously-acquired skills via goal generation
Akakzia, A.; Colas, C.; Oudeyer, P.-Y.; Chetouani, M.; and Sigaud, O. 2020 · 2006
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E.; Erez, T.; and Tassa, Y. 2012 · 2012
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M. G.; Graves, A.; Riedmiller, M.; Fidjeland, A. K.; Ostrovski, G.; et al. 2015 · 2015
Earlier work this paper cites.
Hindsight experience replay
Andrychowicz, M.; Wolski, F.; Ray, A.; Schneider, J.; Fong, R.; Welinder, P.; McGrew, B.; Tobin, J.; Pieter Abbeel, O.; and Zaremba, W. 2017 · 2017
Earlier work this paper cites.
The option-critic architecture
Bacon, P.-L.; Harb, J.; and Precup, D. 2017 · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; and Amodei, D. 2017 · 2017
Cited alongside, same era.
Exploration-exploitation in mdps with options
Fruit, R.; and Lazaric, A. 2017 · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T.; Zhou, A.; Abbeel, P.; and Levine, S. 2018 · 2018
Cited alongside, same era.
Hierarchical imitation and reinforcement learning
Le, H.; Jiang, N.; Agarwal, A.; Dudik, M.; Yue, Y.; and Daumé III, H. 2018 · 2018
Cited alongside, same era.
Improving Safety in Reinforcement Learning Using Model-Based Architectures and Human Intervention
Prakash, B.; et al. 2019 · 2019
Later among the works it cites.
Pyperplan
Alkhazraji, Y.; Frorath, M.; Grützner, M.; Helmert, M.; Liebetraut, T.; Mattmüller, R.; Ortlieb, M.; Seipp, J.; Springenberg, T.; Stahl, P.; and Wülfing, J. 2020 · 2020
Later among the works it cites.
Guiding Safe Reinforcement Learning Policies Using Structured Language Constraints
Prakash, B.; et al. 2020 · 2020
Later among the works it cites.
Interactive Hierarchical Guidance using Language
Prakash, B.; Waytowich, N.; Oates, T.; and Mohsenin, T. 2021 · 2021
Later among the works it cites.
An Energy-Efficient Hardware Accelerator for Hierarchical Deep Reinforcement Learning
Shiri, A.; et al. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep TAMER: Interactive Agent Shaping in High-Dimensional State Spaces
Warnell, G.; Waytowich, N.; Lawhern, V.; and Stone, P. 2018 · 2018
Cited alongside, same era.
Language as an abstraction for hierarchical deep reinforcement learning
Jiang, Y.; Gu, S. S.; Murphy, K. P.; and Finn, C. 2019 · 2019
Cited alongside, same era.
Toward Real-World Implementation of Deep Reinforcement Learning for Vision-Based Autonomous Drone Navigation with Mission
Navardi, M.; Dixit, P.; Manjunath, T.; Waytowich, N. R.; Mohsenin, T.; and Oates, T. 2022 · 2022
Closest in time.
Efficient Language-Guided Reinforcement Learning for Resource Constrained Autonomous Systems
Shiri, A.; Navardi, M.; Manjunath, T.; Waytowich, N. R.; and Mohsenin, T. 2022 · 2022
Closest in time.