Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) agents improve through trial-and-error, but when reward is sparse and the agent cannot discover successful action sequences, learning stagnates.
Efficient training of artificial neural networks for autonomous navigation
D. A. Pomerleau · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Robot learning from demonstration
C. G. Atkeson and S. Schaal · 1997
Earlier work this paper cites.
The MAXQ method for hierarchical reinforcement learning
T. G. Dietterich · 1998
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
R. Parr and S. J. Russell · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
A. Y. Ng, D. Harada, and S. Russell · 1999
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
Programmable reinforcement learning agents
D. Andre · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Ng · 2004
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
N. Chentanez, A. G. Barto, and S. P. Singh · 2005
Earlier work this paper cites.
Concurrent hierarchical reinforcement learning
B. Marthi and C. Guestrin · 2005
Earlier work this paper cites.
Towards a unified theory of state abstraction for mdps
L. Li, T. J. Walsh, and M. L. Littman · 2006
Earlier work this paper cites.
PLOW: A collaborative task learning agent
J. Allen, N. Chambers, G. Ferguson, L. Galescu, H. Jung, M. Swift, and W. Taysom · 2007
Earlier work this paper cites.
Building portable options: Skill transfer in reinforcement learning
G. Konidaris and A. G. Barto · 2007
Earlier work this paper cites.
Using motion primitives in probabilistic sample-based planning for humanoid robots
K. Hauser, T. Bretl, K. Harada, and J. Latombe · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey · 2008
Earlier work this paper cites.
Sikuli: using GUI screenshots for search and automation
T. Yeh, T. Chang, and R. Miller · 2009
Earlier work this paper cites.
Using dimensionality reduction to exploit constraints in reinforcement learning
S. Bitzer, M. Howard, and S. Vijayakumar · 2010
Cited alongside, same era.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and A. Bagnell · 2011
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Cited alongside, same era.
Learning from limited demonstrations
B. Kim, A. massoud Farahmand, J. Pineau, and D. Precup · 2013
Cited alongside, same era.
Guided policy search
S. Levine and V. Koltun · 2013
Cited alongside, same era.
Reinforcement and imitation learning via interactive no-regret learning
S. Ross and J. A. Bagnell · 2014
Deep exploration via bootstrapped DQN
I. Osband, C. Blundell, A. Pritzel, and B. V. Roy · 2016
Later among the works it cites.
Learning to plan for constrained manipulation from demonstrations
M. Phillips, V. Hwang, S. Chitta, and M. Likhachev · 2016
Later among the works it cites.
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, P. Abbeel, and W. Zaremba · 2017
Later among the works it cites.
The option-critic architecture
P. Bacon, J. Harb, and D. Precup · 2017
Later among the works it cites.
End-to-end differentiable adversarial imitation learning
N. Baram, O. Anschel, I. Caspi, and S. Mannor · 2017
Later among the works it cites.
Inductive representation learning on large graphs
W. L. Hamilton, R. Ying, and J. Leskovec · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Amazon Unveils a Listening, Talking, Music-Playing Speaker for Your Home
B. Stone and S. Soper · 2014
Cited alongside, same era.
Reinforcement learning from demonstration through shaping
T. Brys, A. Harutyunyan, H. B. Suay, S. Chernova, M. E. Taylor, and A. Now’e · 2015
Cited alongside, same era.
Building a semantic parser overnight
Y. Wang, J. Berant, and P. Liang · 2015
Cited alongside, same era.
Modular multitask reinforcement learning with policy sketches
J. Andreas, D. Klein, and S. Levine · 2016
Cited alongside, same era.
Hierarchical relative entropy policy search
C. Daniel, G. Neumann, O. Kroemer, and J. Peters · 2016
Cited alongside, same era.
Erica: Interaction mining mobile apps
B. Deka, Z. Huang, and R. Kumar · 2016
Cited alongside, same era.
Later among the works it cites.
Deep reward shaping from demonstrations
A. Hussein, E. Elyan, M. M. Gaber, and C. Jayne · 2017
Later among the works it cites.
Semi-supervised classification with graph convolutional networks
T. N. Kipf and M. Welling · 2017
Later among the works it cites.
Overcoming exploration in reinforcement learning with demonstrations
A. Nair, B. McGrew, M. Andrychowicz, W. Zaremba, and P. Abbeel · 2017
Later among the works it cites.
C-learn: Learning geometric constraints from demonstrations for multi-step manipulation in shared autonomy
C. Perez-D’Arpino and J. A. Shah · 2017
Later among the works it cites.
Column networks for collective classification
T. Pham, T. Tran, D. Phung, and S. Venkatesh · 2017
Later among the works it cites.
M. Roderick, C. Grimm, and S. Tellex · 2017
Later among the works it cites.
World of bits: An open-domain platform for web-based agents
T. Shi, A. Karpathy, L. Fan, J. Hernandez, and P. Liang · 2017
Later among the works it cites.
Deeply aggrevated: Differentiable imitation learning for sequential prediction
W. Sun, A. Venkatraman, G. J. Gordon, B. Boots, and J. A. Bagnell · 2017
Later among the works it cites.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
M. Vecerik, T. Hester, J. Scholz, F. Wang, O. Pietquin, B. Piot, N. Heess, T. Rothorl, T. Lampe, and M. Riedmiller · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
T. Weber, S. Racanière, D. P. Reichert, L. Buesing, A. Guez, D. J. Rezende, A. P. Badia, O. Vinyals, N. Heess, Y. Li, et al · 2017
Later among the works it cites.
Deep Q-learning from demonstrations
T. Hester, M. Vecerik, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, A. Sendonaris, G. Dulac-Arnold, I. Osband, J. Agapiou, J. Z. Leibo, and A. Gruslys · 2018
Closest in time.