Fetching the paper…
Reading the bibliography…
In this paper, hypernetworks are trained to generate behaviors across a range of unseen task conditions, via a novel TD-based training objective and data from a set of near-optimal RL solutions for training tasks.
Reinforcement Learning Upside Down: Don’t Predict Rewards–Just Map Them to Actions
Schmidhuber, J. 2019 · 1912
Earlier work this paper cites.
Training agents using upside-down reinforcement learning
Srivastava, R. K.; Shyam, P.; Mutz, F.; Jaśkowski, W.; and Schmidhuber, J. 2019 · 1912
Earlier work this paper cites.
Dynamic programming
Bellman, R. 1966 · 1966
Earlier work this paper cites.
Harb, J.; Schaul, T.; Precup, D.; and Bacon, P.-L. 2020 · 2002
Earlier work this paper cites.
Schroecker, Y.; and Isbell, C. 2020 · 2002
Earlier work this paper cites.
Deep multi-agent reinforcement learning for decentralized continuous cooperative control
de Witt, C. S.; Peng, B.; Kamienny, P.-A.; Torr, P.; Böhmer, W.; and Whiteson, S. 2020 · 2003
Earlier work this paper cites.
Supervised actor-critic reinforcement learning
Rosenstein, M. T.; Barto, A. G.; Si, J.; Barto, A.; Powell, W.; and Wunsch, D. 2004 · 2004
Earlier work this paper cites.
Ai-qmix: Attention and imagination for dynamic multi-agent reinforcement learning
Iqbal, S.; de Witt, C. A. S.; Peng, B.; Böhmer, W.; Whiteson, S.; and Sha, F. 2020 · 2006
Earlier work this paper cites.
learn2learn: A library for meta-learning research
Arnold, S. M.; Mahajan, P.; Datta, D.; Bunner, I.; and Zarkias, K. S. 2020 · 2008
Earlier work this paper cites.
Transfer Learning for Reinforcement Learning Domains: A Survey
Taylor, M. E.; and Stone, P. 2009 · 2009
Earlier work this paper cites.
Reward design via online gradient ascent
Sorg, J.; Lewis, R. L.; and Singh, S. 2010 · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S.; Gordon, G.; and Bagnell, D. 2011 · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E.; Erez, T.; and Tassa, Y. 2012 · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Graves, A.; Antonoglou, I.; Wierstra, D.; and Riedmiller, M. 2013 · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L. 2014 · 2014
Earlier work this paper cites.
An invitation to imitation
Bagnell, J. A. 2015 · 2015
Earlier work this paper cites.
Contextual markov decision processes
Hallak, A.; Di Castro, D.; and Mannor, S. 2015 · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P.; Hunt, J. J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; and Wierstra, D. 2015 · 2015
Earlier work this paper cites.
Universal value function approximators
Schaul, T.; Horgan, D.; Gregor, K.; and Silver, D. 2015 · 2015
Earlier work this paper cites.
Deep learning in neural networks: An overview
Schmidhuber, J. 2015 · 2015
Cited alongside, same era.
Ha, D.; Dai, A.; and Le, Q. V. 2016 · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Ho, J.; and Ermon, S. 2016 · 2016
Cited alongside, same era.
Hindsight experience replay
Andrychowicz, M.; Wolski, F.; Ray, A.; Schneider, J.; Fong, R.; Welinder, P.; McGrew, B.; Tobin, J.; Pieter Abbeel, O.; and Zaremba, W. 2017 · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C.; Abbeel, P.; and Levine, S. 2017 · 2017
Cited alongside, same era.
On learning intrinsic rewards for policy gradient methods
Zheng, Z.; Oh, J.; and Singh, S. 2018 · 2018
Later among the works it cites.
Rudder: Return decomposition for delayed rewards
Arjona-Medina, J. A.; Gillhofer, M.; Widrich, M.; Unterthiner, T.; Brandstetter, J.; and Hochreiter, S. 2019 · 2019
Later among the works it cites.
Hindsight credit assignment
Harutyunyan, A.; Dabney, W.; Mesnard, T.; Gheshlaghi Azar, M.; Piot, B.; Heess, N.; van Hasselt, H. P.; Wayne, G.; Singh, S.; Precup, D.; et al. 2019 · 2019
Later among the works it cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Rakelly, K.; Zhou, A.; Finn, C.; Levine, S.; and Quillen, D. 2019 · 2019
Later among the works it cites.
Continual learning with hypernetworks
von Oswald, J.; Henning, C.; Grewe, B. F.; and Sacramento, J. 2019 · 2019
Later among the works it cites.
Fast reinforcement learning with generalized policy updates
Barreto, A.; Hou, S.; Borsa, D.; Silver, D.; and Precup, D. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Contextual decision processes with low bellman rank are pac-learnable
Jiang, N.; Krishnamurthy, A.; Agarwal, A.; Langford, J.; and Schapire, R. E. 2017 · 2017
Cited alongside, same era.
Krueger, D.; Huang, C.-W.; Islam, R.; Turner, R.; Lacoste, A.; and Courville, A. 2017 · 2017
Cited alongside, same era.
Model-based contextual policy search for data-efficient generalization of robot skills
Kupcsik, A.; Deisenroth, M. P.; Peters, J.; Loh, A. P.; Vadakkepat, P.; and Neumann, G. 2017 · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Paszke, A.; Gross, S.; Chintala, S.; Chanan, G.; Yang, E.; DeVito, Z.; Lin, Z.; Desmaison, A.; Antiga, L.; and Lerer, A. 2017 · 2017
Cited alongside, same era.
Rauber, P.; Ummadisingu, A.; Mutz, F.; and Schmidhuber, J. 2017 · 2017
Cited alongside, same era.
Distral: Robust multitask reinforcement learning
Teh, Y.; Bapst, V.; Czarnecki, W. M.; Quan, J.; Kirkpatrick, J.; Hadsell, R.; Heess, N.; and Pascanu, R. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
On the modularity of hypernetworks
Galanti, T.; and Wolf, L. 2020 · 2020
Later among the works it cites.
Teacher algorithms for curriculum learning of Deep RL in continuously parameterized environments
Portelas, R.; Colas, C.; Hofmann, K.; and Oudeyer, P.-Y. 2020 · 2020
Later among the works it cites.
Meta-gradient reinforcement learning with an objective discovered online
Xu, Z.; van Hasselt, H. P.; Hessel, M.; Oh, J.; Singh, S.; and Silver, D. 2020 · 2020
Later among the works it cites.
Meta-learning via hypernetworks
Zhao, D.; von Oswald, J.; Kobayashi, S.; Sacramento, J.; and Grewe, B. F. 2020 · 2020
Later among the works it cites.
What can learned intrinsic rewards capture?
Zheng, Z.; Oh, J.; Hessel, M.; Xu, Z.; Kroiss, M.; Van Hasselt, H.; Silver, D.; and Singh, S. 2020 · 2020
Later among the works it cites.
Learning implicit credit assignment for cooperative multi-agent reinforcement learning
Zhou, M.; Liu, Z.; Sui, P.; Li, Y.; and Chung, Y. Y. 2020 · 2020
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L.; Lu, K.; Rajeswaran, A.; Lee, K.; Grover, A.; Laskin, M.; Abbeel, P.; Srinivas, A.; and Mordatch, I. 2021 · 2021
Later among the works it cites.
Continual model-based reinforcement learning with hypernetworks
Huang, Y.; Xie, K.; Bharadhwaj, H.; and Shkurti, F. 2021 · 2021
Later among the works it cites.
Randomized Entity-wise Factorization for Multi-Agent Reinforcement Learning
Iqbal, S.; De Witt, C. A. S.; Peng, B.; Böhmer, W.; Whiteson, S.; and Sha, F. 2021 · 2021
Later among the works it cites.
Offline reinforcement learning as one big sequence modeling problem
Janner, M.; Li, Q.; and Levine, S. 2021 · 2021
Later among the works it cites.
Recomposing the reinforcement learning building blocks with hypernetworks
Sarafian, E.; Keynan, S.; and Kraus, S. 2021 · 2021
Later among the works it cites.
Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble
Lee, S.; Seo, Y.; Lee, K.; Abbeel, P.; and Shin, J. 2022 · 2022
Closest in time.
Operator Deep Q-Learning: Zero-Shot Reward Transferring in Reinforcement Learning
Tang, Z.; Feng, Y.; and Liu, Q. 2022 · 2022
Closest in time.