Fetching the paper…
Reading the bibliography…
One of the main challenges in reinforcement learning is solving tasks with sparse reward.
Über ein paradoxon aus der verkehrsplanung
Braess, D · 1968
Earlier work this paper cites.
Algebraic connectivity of graphs
Fiedler, M · 1973
Earlier work this paper cites.
Bounds on the cover time
Broder, A. Z. and Karlin, A. R · 1989
Earlier work this paper cites.
A heuristic approach to the discovery of macro-operators
Iba, G. A · 1989
Earlier work this paper cites.
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
Interlacing eigenvalues and graphs
Haemers, W. H · 1995
Earlier work this paper cites.
Spectral graph theory
Chung, F. R · 1996
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
Parr, R. and Russell, S · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Sutton, R., , Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the MAXQ value function decomposition
Dietterich, T. G · 2000
Earlier work this paper cites.
Automatic discovery of subgoals in reinforcement learning using diverse density
McGovern, A. and Barto, A. G · 2001
Earlier work this paper cites.
Ace: A fast multiscale eigenvectors computation for drawing huge graphs
Koren, Y., Carmel, L., and Harel, D · 2002
Cited alongside, same era.
Q-cut - dynamic discovery of sub-goals in reinforcement learning
Menache, I., Mannor, S., and Shimkin, N · 2002
Cited alongside, same era.
Learning options in reinforcement learning
Stolle, M. and Precup, D · 2002
Cited alongside, same era.
On spectral graph drawing
Koren, Y · 2003
Cited alongside, same era.
Using relative novelty to identify useful temporal abstractions in reinforcement learning
Şimşek, Ö. and Barto, A · 2004
Cited alongside, same era.
On a paradox of traffic planning
Braess, D., Nagurney, A., and Wakolbinger, T · 2005
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P · 2010
Later among the works it cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M · 2014
Later among the works it cites.
Learning purposeful behaviour in the absence of rewards
Machado, M. C. and Bowling, M · 2016
Later among the works it cites.
Adaptive skills adaptive partitions (ASAP)
Mankowitz, D. J., Mann, T. A., and Mannor, S · 2016
Later among the works it cites.
The option-critic architecture
Bacon, P.-L., Harb, J., and Precup, D · 2017
Later among the works it cites.
When waiting is not an option: Learning options with a deliberation cost
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Identifying useful subgoals in reinforcement learning by local graph partitioning
Simsek, O., Wolfe, A., and Barto, A · 2005
Cited alongside, same era.
Growing well-connected graphs
Ghosh, A. and Boyd, S · 2006
Cited alongside, same era.
Logarithmic online regret bounds for undiscounted reinforcement learning
Ortner, P. and Auer, R · 2007
Cited alongside, same era.
Maximum algebraic connectivity augmentation is np-hard
Mosk-Aoyama, D · 2008
Cited alongside, same era.
Skill discovery in continuous reinforcement learning domains using skill chaining
Konidaris, G. and Barto, A · 2009
Cited alongside, same era.
Skill characterization based on betweenness
Şimşek, Ö. and Barto, A. G · 2009
Cited alongside, same era.
Harb, J., Bacon, P.-L., Klissarov, M., and Precup, D · 2017
Later among the works it cites.
Continual curiosity-driven skill acquisition from high-dimensional video inputs for humanoid robots
Kompella, V. R., Stollenga, M., Luciw, M., and Schmidhuber, J · 2017
Later among the works it cites.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2018
Later among the works it cites.
Finding options that minimize planning time
Jinnai, Y., Abel, D., Littman, M., and Konidaris, G · 2018
Later among the works it cites.
Visual reinforcement learning with imagined goals
Nair, A. V., Pong, V., Dalal, M., Bahl, S., Lin, S., and Levine, S · 2018
Later among the works it cites.
Learning by playing solving sparse reward tasks from scratch
Riedmiller, M., Hafner, R., Lampe, T., Neunert, M., Degrave, J., van de Wiele, T., Mnih, V., Heess, N., and Springenberg, J. T · 2018
Later among the works it cites.