Fetching the paper…
Reading the bibliography…
As reinforcement learning agents are tasked with solving more challenging and diverse tasks, the ability to incorporate prior knowledge into the learning system and to exploit reusable structure in solution space is likely to become increasingly important.
Feudal reinforcement learning
Dayan, P. and Hinton, G. E · 1993
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
Parr, R. and Russell, S. J · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
An auxiliary variational method
Agakov, F. V. and Barber, D · 2004
Earlier work this paper cites.
Linearly-solvable markov decision problems
Todorov, E · 2007
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
Toussaint, M · 2009
Earlier work this paper cites.
Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy
Ziebart, B. D · 2010
Earlier work this paper cites.
Information theory of decisions and actions
Tishby, N. and Polani, D · 2011
Earlier work this paper cites.
Optimal control as a graphical model inference problem
Kappen, H. J., Gómez, V., and Opper, M · 2012
Earlier work this paper cites.
On stochastic optimal control and reinforcement learning by approximate inference
Rawlik, K., Toussaint, M., and Vijayakumar, S · 2012
Earlier work this paper cites.
Trading value and information in mdps
Rubin, J., Shamir, O., and Tishby, N · 2012
Earlier work this paper cites.
An information-theoretic approach to curiosity-driven reinforcement learning
Still, S. and Precup, D · 2012
Earlier work this paper cites.
Thermodynamics as a theory of decision-making with information-processing costs
Ortega, P. A. and Braun, D. A · 2013
Earlier work this paper cites.
Markov Chain Monte Carlo and Variational Inference: Bridging the Gap
Salimans, T., Kingma, D. P., and Welling, M · 2014
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
Heess, N., Wayne, G., Silver, D., Lillicrap, T., Erez, T., and Tassa, Y · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Moritz, P., Jordan, M., and Abbeel, P · 2015
Earlier work this paper cites.
Taming the noise in reinforcement learning via soft updates
Fox, R., Pakman, A., and Tishby, N · 2016
Earlier work this paper cites.
Learning and transfer of modulated locomotor controllers
Heess, N., Wayne, G., Tassa, Y., Lillicrap, T., Riedmiller, M., and Silver, D · 2016
Earlier work this paper cites.
Composing graphical models with neural networks for structured representations and fast inference
Johnson, M., Duvenaud, D. K., Wiltschko, A., Adams, R. P., and Datta, S. R · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M · 2016
Cited alongside, same era.
Actor-mimic: Deep multitask and transfer reinforcement learning
Parisotto, E., Ba, J. L., and Salakhutdinov, R · 2016
Cited alongside, same era.
Policy distillation
Rusu, A. A., Colmenarejo, S. G., Gulcehre, C., Desjardins, G., Kirkpatrick, J., Pascanu, R., Mnih, V., Kavukcuoglu, K., and Hadsell, R · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
Mix & match agent curricula for reinforcement learning
Czarnecki, W., Jayakumar, S., Jaderberg, M., Hasenclever, L., Teh, Y. W., Heess, N., Osindero, S., and Pascanu, R · 2018
Later among the works it cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al · 2018
Later among the works it cites.
Meta learning shared hierarchies
Frans, K., Ho, J., Chen, X., Abbeel, P., and Schulman, J · 2018
Later among the works it cites.
Divide-and-conquer reinforcement learning
Ghosh, D., Singh, A., Rajeswaran, A., Kumar, V., and Levine, S · 2018
Later among the works it cites.
Learning an embedding space for transferable robot skills
Hausman, K., Springenberg, J. T., Wang, Z., Heess, N., and Riedmiller, M · 2018
Later among the works it cites.
Data-efficient hierarchical reinforcement learning
Nachum, O., Gu, S. S., Lee, H., and Levine, S · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Cited alongside, same era.
The option-critic architecture
Bacon, P.-L., Harb, J., and Precup, D · 2017
Cited alongside, same era.
Stochastic neural networks for hierarchical reinforcement learning
Florensa, C., Duan, Y., and Abbeel, P · 2017
Cited alongside, same era.
Multi-level discovery of deep options
Fox, R., Krishnan, S., Stoica, I., and Goldberg, K · 2017
Cited alongside, same era.
Variational intrinsic control
Gregor, K., Rezende, D. J., and Wierstra, D · 2017
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Emergence of locomotion behaviours in rich environments
Heess, N., Tirumala, D., Sriram, S., Lemmon, J., Merel, J., Wayne, G., Tassa, Y., Erez, T., Wang, Z., Eslami, A., Riedmiller, M., et al · 2017
Cited alongside, same era.
Later among the works it cites.
Openai five
OpenAI · 2018
Later among the works it cites.
Learning dexterous in-hand manipulation
OpenAI, Andrychowicz, M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al · 2018
Later among the works it cites.
Learning by playing solving sparse reward tasks from scratch
Riedmiller, M., Hafner, R., Lampe, T., Neunert, M., Degrave, J., van de Wiele, T., Mnih, V., Heess, N., and Springenberg, J. T · 2018
Later among the works it cites.
Kickstarting deep reinforcement learning
Schmitt, S., Hudson, J. J., Zidek, A., Osindero, S., Doersch, C., Czarnecki, W. M., Leibo, J. Z., Kuttler, H., Zisserman, A., Simonyan, K., et al · 2018
Later among the works it cites.
Learning to share and hide intentions using information regularization
Strouse, D., Kleiman-Weiner, M., Tenenbaum, J., Botvinick, M., and Schwab, D. J · 2018
Later among the works it cites.
Transferring task goals via hierarchical reinforcement learning, 2018
Xie, S., Galashov, A., Liu, S., Hou, S., Pascanu, R., Heess, N., and Teh, Y. W · 2018
Later among the works it cites.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2019
Closest in time.
Information asymmetry in KL-regularized RL
Galashov, A., Jayakumar, S., Hasenclever, L., Tirumala, D., Schwarz, J., Desjardins, G., Czarnecki, W. M., Teh, Y. W., Pascanu, R., and Heess, N · 2019
Closest in time.
Transfer and exploration via the information bottleneck
Goyal, A., Islam, R., Strouse, D., Ahmed, Z., Larochelle, H., Botvinick, M., Levine, S., and Bengio, Y · 2019
Closest in time.
Soft q-learning with mutual-information regularization
Grau-Moya, J., Leibfried, F., and Vrancx, P · 2019
Closest in time.
Composing complex skills by learning transition policies with proximity reward induction
Lee, Y., Sun, S.-H., Somasundaram, S., Hu, E., and Lim, J. J · 2019
Closest in time.
Neural probabilistic motor primitives for humanoid control
Merel, J., Hasenclever, L., Galashov, A., Ahuja, A., Pham, V., Wayne, G., Teh, Y. W., and Heess, N · 2019
Closest in time.
Near-optimal representation learning for hierarchical reinforcement learning
Nachum, O., Gu, S., Lee, H., and Levine, S · 2019
Closest in time.