Fetching the paper…
Reading the bibliography…
Humans can abstract prior knowledge from very little data and use it to boost skill learning.
SQIL: Imitation Learning via Regularized Behavioral Cloning URL
Reddy, S.; Dragan, A. D.; and Levine, S. 2019 · 1905
Earlier work this paper cites.
Construction of Macro Actions for Deep Reinforcement Learning
Chang, Y.-H.; Chang, K.-Y.; Kuo, H.; and Lee, C.-Y. 2019 · 1908
Earlier work this paper cites.
STRIPS: A new approach to the application of theorem proving to problem solving
Fikes, R. E.; and Nilsson, N. J. 1971 · 1971
Earlier work this paper cites.
Learning and Executing Generalized Robot Plans
Fikes, R. E.; Hart, P. E.; and Nilsson, N. J. 1972 · 1972
Earlier work this paper cites.
The Role of Preprocessing in Problem Solving Systems: “An Ounce of Reflection is Worth a Pound of Backtracking”
Dawson, C.; and Siklossy, L. 1977 · 1977
Earlier work this paper cites.
Selectively Generalizing Plans for Problem-Solving
Minton, S. 1985 · 1985
Earlier work this paper cites.
Efficient exploration in reinforcement learning
Thrun, S. B. 1992 · 1992
Earlier work this paper cites.
Behavioural cloning in control of a dynamic system
Esmaili, N.; Sammut, C.; and Shirazi, G. M. 1995 · 1995
Earlier work this paper cites.
Identifying Hierarchical Structure in Sequences: A linear-time algorithm
Nevill-Manning, C. G.; and Witten, I. H. 1997 · 1997
Earlier work this paper cites.
Macro-actions in reinforcement learning: An empirical analysis
McGovern, A.; and Sutton, R. S. 1998 · 1998
Earlier work this paper cites.
Learning Macro-Actions in Reinforcement Learning
Randlov, J. 1999 · 1999
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Sutton, R. S.; Precup, D.; and Singh, S. 1999 · 1999
Earlier work this paper cites.
PolicyBlocks: An Algorithm for Creating Useful Macro-Actions in Reinforcement Learning
Pickett, M.; and Barto, A. G. 2002 · 2002
Earlier work this paper cites.
Learning options in reinforcement learning
Stolle, M.; and Precup, D. 2002 · 2002
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
Barto, A. G.; and Mahadevan, S. 2003 · 2003
Earlier work this paper cites.
A survey of robot learning from demonstration
Argall, B. D.; Chernova, S.; Veloso, M.; and Browning, B. 2009 · 2009
Earlier work this paper cites.
Levenshtein Distance: Information Theory, Computer Science, String (Computer Science), String Metric, Damerau?Levenshtein Distance, Spell Checker, Hamming Distance
Miller, F. P.; Vandome, A. F.; and McBrewster, J. 2009 · 2009
Cited alongside, same era.
No-Regret Reductions for Imitation Learning and Structured Prediction
Ross, S.; Gordon, G. J.; and Bagnell, J. A. 2010 · 2010
Cited alongside, same era.
Bayesian policy search with policy priors
Wingate, D.; Goodman, N. D.; Roy, D. M.; Kaelbling, L. P.; and Tenenbaum, J. B. 2011 · 2011
Cited alongside, same era.
The Arcade Learning Environment: An Evaluation Platform for General Agents
Bellemare, M. G.; Naddaf, Y.; Veness, J.; and Bowling, M. 2012 · 2012
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Kingma, D. P.; and Ba, J. 2015 · 2015
Cited alongside, same era.
Curiosity-driven Exploration by Self-supervised Prediction
Pathak, D.; Agrawal, P.; Efros, A. A.; and Darrell, T. 2017 · 2017
Later among the works it cites.
Learning to repeat: Fine grained action repetition for deep reinforcement learning
Sharma, S.; Lakshminarayanan, A. S.; and Ravindran, B. 2017 · 2017
Later among the works it cites.
Human learning in Atari
Tsividis, P. A.; Pouncy, T.; Xu, J. L.; Tenenbaum, J. B.; and Gershman, S. J. 2017 · 2017
Later among the works it cites.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
Vecerik, M.; Hester, T.; Scholz, J.; Wang, F.; Pietquin, O.; Piot, B.; Heess, N.; Rothörl, T.; Lampe, T.; and Riedmiller, M. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Human-level control through deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M. G.; Graves, A.; Riedmiller, M.; Fidjeland, A. K.; Ostrovski, G.; Petersen, S.; Beattie, C.; Sadik, A.; Antonoglou, I.; King, H.; Kumaran, D.; Wierstra, D.; Legg, S.; and Hassabis, D. 2015 · 2015
Cited alongside, same era.
The Option-Critic Architecture
Bacon, P.; Harb, J.; and Precup, D. 2016 · 2016
Cited alongside, same era.
Brockman, G.; Cheung, V.; Pettersson, L.; Schneider, J.; Schulman, J.; Tang, J.; and Zaremba, W. 2016 · 2016
Cited alongside, same era.
Deep Reinforcement Learning With Macro-Actions
Durugkar, I. P.; Rosenbaum, C.; Dernbach, S.; and Mahadevan, S. 2016 · 2016
Cited alongside, same era.
Generative Adversarial Imitation Learning
Ho, J.; and Ermon, S. 2016 · 2016
Cited alongside, same era.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Kulkarni, T. D.; Narasimhan, K.; Saeedi, A.; and Tenenbaum, J. 2016 · 2016
Cited alongside, same era.
Asynchronous Methods for Deep Reinforcement Learning
Mnih, V.; Badia, A. P.; Mirza, M.; Graves, A.; Lillicrap, T. P.; Harley, T.; Silver, D.; and Kavukcuoglu, K. 2016 · 2016
Cited alongside, same era.
Cobbe, K.; Klimov, O.; Hesse, C.; Kim, T.; and Schulman, J. 2018 · 2018
Later among the works it cites.
IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Espeholt, L.; Soyer, H.; Munos, R.; Simonyan, K.; Mnih, V.; Ward, T.; Doron, Y.; Firoiu, V.; Harley, T.; Dunning, I.; Legg, S.; and Kavukcuoglu, K. 2018 · 2018
Later among the works it cites.
Deep q-learning from demonstrations
Hester, T.; Vecerik, M.; Pietquin, O.; Lanctot, M.; Schaul, T.; Piot, B.; Horgan, D.; Quan, J.; Sendonaris, A.; Osband, I.; et al. 2018 · 2018
Later among the works it cites.
Policy Optimization with Demonstrations
Kang, B.; Jie, Z.; and Feng, J. 2018 · 2018
Later among the works it cites.
Hierarchical Imitation and Reinforcement Learning
Le, H. M.; Jiang, N.; Agarwal, A.; Dudík, M.; Yue, Y.; and III, H. D. 2018 · 2018
Later among the works it cites.
Overcoming exploration in reinforcement learning with demonstrations
Nair, A.; McGrew, B.; Andrychowicz, M.; Zaremba, W.; and Abbeel, P. 2018 · 2018
Later among the works it cites.
Learning Abstract Options
Riemer, M.; Liu, M.; and Tesauro, G. 2018 · 2018
Later among the works it cites.
Learning Montezuma’s Revenge from a Single Demonstration
Salimans, T.; and Chen, R. 2018 · 2018
Later among the works it cites.
Reinforcement Learning with Structured Hierarchical Grammar Representations of Actions
Christodoulou, P.; Lange, R. T.; Shafti, A.; and Faisal, A. A. 2019 · 2019
Later among the works it cites.
A Compression-Inspired Framework for Macro Discovery
Garcia, F. M.; da Silva, B. C.; and Thomas, P. S. 2019 · 2019
Later among the works it cites.
CompILE: Compositional Imitation Learning and Execution
Kipf, T.; Li, Y.; Dai, H.; Zambaldi, V.; Sanchez-Gonzalez, A.; Grefenstette, E.; Kohli, P.; and Battaglia, P. 2019 · 2019
Later among the works it cites.
Discovering Motor Programs by Recomposing Demonstrations
Shankar, T.; Tulsiani, S.; Pinto, L.; and Gupta, A. 2020 · 2020
Closest in time.