Fetching the paper…
Reading the bibliography…
Reinforcement learning has the potential to automate the acquisition of behavior in complex settings, but in order for it to be successfully deployed, a number of practical challenges must be addressed.
Marktform und Gleichgewicht
H. Von Stackelberg · 1934
Earlier work this paper cites.
Curious model-building control systems
J. Schmidhuber · 1991
Earlier work this paper cites.
Feudal reinforcement learning
P. Dayan and G. E. Hinton · 1993
Earlier work this paper cites.
Temporal difference learning and td-gammon
G. Tesauro · 1995
Earlier work this paper cites.
Hq-learning
M. Wiering and J. Schmidhuber · 1997
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
R. Parr and S. J. Russell · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
T. G. Dietterich · 2000
Earlier work this paper cites.
Temporal abstraction in reinforcement learning
D. Precup · 2001
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
M. Kearns and S. Singh · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
S. M. Kakade et al · 2003
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
N. Chentanez, A. G. Barto, and S. P. Singh · 2005
Earlier work this paper cites.
Empowerment: A universal agent-centric measure of control
A. S. Klyubin, D. Polani, and C. L. Nehaniv · 2005
Earlier work this paper cites.
An overview of bilevel optimization
B. Colson, P. Marcotte, and G. Savard · 2007
Earlier work this paper cites.
Active learning of inverse models with intrinsically motivated goal exploration in robots
A. Baranes and P.-Y. Oudeyer · 2013
Earlier work this paper cites.
Intrinsic motivation and reinforcement learning
A. G. Barto · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
Empowerment–an introduction
C. Salge, C. Glackin, and D. Polani · 2014
Earlier work this paper cites.
Learning compound multi-step controllers under unknown dynamics
W. Han, S. Levine, and P. Abbeel · 2015
Cited alongside, same era.
Variational information maximisation for intrinsically motivated reinforcement learning
S. Mohamed and D. J. Rezende · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Cited alongside, same era.
K. Gregor, D. J. Rezende, and D. Wierstra · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
H. Van Hasselt, A. Guez, and D. Silver · 2016
Reset-free trial-and-error learning for robot damage recovery
K. Chatzilygeroudis, V. Vassiliades, and J.-B. Mouret · 2018
Later among the works it cites.
Diversity is all you need: Learning skills without a reward function
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine · 2018
Later among the works it cites.
Variational inverse control with events: A general framework for data-driven reward definition
J. Fu, A. Singh, D. Ghosh, L. Yang, and S. Levine · 2018
Later among the works it cites.
Composable deep reinforcement learning for robotic manipulation
T. Haarnoja, V. Pong, A. Zhou, M. Dalal, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
Data-efficient hierarchical reinforcement learning
O. Nachum, S. S. Gu, H. Lee, and S. Levine · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The option-critic architecture
P.-L. Bacon, J. Harb, and D. Precup · 2017
Cited alongside, same era.
Minimal criterion coevolution: a new approach to open-ended search
J. C. Brant and K. O. Stanley · 2017
Cited alongside, same era.
Leave no trace: Learning to reset for safe and autonomous reinforcement learning
B. Eysenbach, S. Gu, J. Ibarz, and S. Levine · 2017
Cited alongside, same era.
Stochastic neural networks for hierarchical reinforcement learning
C. Florensa, Y. Duan, and P. Abbeel · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Cited alongside, same era.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
A. Rajeswaran, V. Kumar, A. Gupta, G. Vezzani, J. Schulman, E. Todorov, and S. Levine · 2017
Cited alongside, same era.
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, et al · 2018
Later among the works it cites.
Emergent tool use from multi-agent autocurricula
B. Baker, I. Kanitscheider, T. Markov, Y. Wu, G. Powell, B. McGrew, and I. Mordatch · 2019
Later among the works it cites.
Convergence of learning dynamics in stackelberg games
T. Fiez, B. Chasnov, and L. J. Ratliff · 2019
Later among the works it cites.
Self-supervised learning of image embedding for continuous control
C. Florensa, J. Degrave, N. Heess, J. T. Springenberg, and M. Riedmiller · 2019
Later among the works it cites.
Efficient exploration via state marginal matching
L. Lee, B. Eysenbach, E. Parisotto, E. Xing, S. Levine, and R. Salakhutdinov · 2019
Later among the works it cites.
Skew-fit: State-covering self-supervised reinforcement learning
V. H. Pong, M. Dalal, S. Lin, A. Nair, S. Bahl, and S. Levine · 2019
Later among the works it cites.
Automated curricula through setter-solver interactions
S. Racaniere, A. K. Lampinen, A. Santoro, D. P. Reichert, V. Firoiu, and T. P. Lillicrap · 2019
Later among the works it cites.
Dynamics-aware unsupervised discovery of skills
A. Sharma, S. Gu, S. Levine, V. Kumar, and K. Hausman · 2019
Later among the works it cites.
R. Wang, J. Lehman, J. Clune, and K. O. Stanley · 2019
Later among the works it cites.
Dexterous manipulation with deep reinforcement learning: Efficient, general, and low-cost
H. Zhu, A. Gupta, A. Rajeswaran, S. Levine, and V. Kumar · 2019
Later among the works it cites.
D4RL: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Closest in time.
The ingredients of real world robotic reinforcement learning
H. Zhu, J. Yu, A. Gupta, D. Shah, K. Hartikainen, A. Singh, V. Kumar, and S. Levine · 2020
Closest in time.