Fetching the paper…
Reading the bibliography…
Deep Reinforcement Learning (RL) is mainly studied in a setting where the training and the testing environments are similar.
Improving exploration in soft-actor-critic with normalizing flows policies
P. N. Ward, A. Smofsky, and A. J. Bose · 1906
Earlier work this paper cites.
P. Kamienny, M. Pirotta, A. Lazaric, T. Lavril, N. Usunier, and L. Denoyer · 2005
Earlier work this paper cites.
Multi-task reinforcement learning: A hierarchical bayesian approach
A. Wilson, A. Fern, S. Ray, and P. Tadepalli · 2007
Earlier work this paper cites.
Robust deep reinforcement learning through adversarial loss
T. P. Oikarinen, T. Weng, and L. Daniel · 2008
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
M. E. Taylor and P. Stone · 2009
Earlier work this paper cites.
One solution is not all you need: Few-shot extrapolation via structured maxent RL
S. Kumar, A. Kumar, S. Levine, and C. Finn · 2010
Earlier work this paper cites.
Transfer in reinforcement learning: a framework and a survey
A. Lazaric · 2012
Earlier work this paper cites.
Openai gym, 2016
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Rl$ˆ2$: Fast reinforcement learning via slow reinforcement learning
Y. Duan, J. Schulman, X. Chen, P. L. Bartlett, I. Sutskever, and P. Abbeel · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Harley, T. P. Lillicrap, D. Silver, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
C. Finn, P. Abbeel, and S. Levine · 2017
Earlier work this paper cites.
Benchmark environments for multitask learning in continuous domains
P. Henderson, W.-D. Chang, F. Shkurti, J. Hansen, D. Meger, and G. Dudek · 2017
Earlier work this paper cites.
Sim-to-Real Transfer of Robotic Control with Dynamics Randomization
X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, T. P. Lillicrap, K. Simonyan, and D. Hassabis · 2017
Cited alongside, same era.
Distral: Robust multitask reinforcement learning
Y. W. Teh, V. Bapst, W. M. Czarnecki, J. Quan, J. Kirkpatrick, R. Hadsell, N. Heess, and R. Pascanu · 2017
Cited alongside, same era.
Meta reinforcement learning as task inference
J. Humplik, A. Galashov, L. Hasenclever, P. A. Ortega, Y. Whye Teh, and N. Heess · 2019
Later among the works it cites.
Explaining landscape connectivity of low-cost solutions for multilayer nets
R. Kuditipudi, X. Wang, H. Lee, Y. Zhang, Z. Li, W. Hu, R. Ge, and S. Arora · 2019
Later among the works it cites.
One solution is not all you need: Few-shot extrapolation via structured maxent RL
S. Kumar, A. Kumar, S. Levine, and C. Finn · 2020
Later among the works it cites.
Loss surface simplexes for mode connecting volumes and fast ensembling
G. W. Benton, W. Maddox, S. Lotfi, and A. G. Wilson · 2021
Closest in time.
Salina: Sequential learning of agents
L. Denoyer, A. de la Fuente, S. Duong, J.-B. Gaya, P.-A. Kamienny, and D. H. Thompson · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Chevalier-Boisvert, L. Willems, and S. Pal · 2018
Cited alongside, same era.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures, 2018
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, S. Legg, and K. Kavukcuoglu · 2018
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine · 2018
Cited alongside, same era.
Loss surfaces, mode connectivity, and fast ensembling of dnns
T. Garipov, P. Izmailov, D. Podoprikhin, D. P. Vetrov, and A. G. Wilson · 2018
Cited alongside, same era.
Learning an embedding space for transferable robot skills
K. Hausman, J. T. Springenberg, Z. Wang, N. Heess, and M. A. Riedmiller · 2018
Cited alongside, same era.
Assessing generalization in deep reinforcement learning, 2018
C. Packer, K. Gao, J. Kos, P. Krähenbühl, V. Koltun, and D. Song · 2018
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation, 2018
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel · 2018
Cited alongside, same era.
C. D. Freeman, E. Frey, A. Raichuk, S. Girgin, I. Mordatch, and O. Bachem · 2021
Closest in time.
Decoupling exploration and exploitation for meta-reinforcement learning without sacrifices
E. Z. Liu, A. Raghunathan, P. Liang, and C. Finn · 2021
Closest in time.
Linear mode connectivity in multitask and continual learning
S. Mirzadeh, M. Farajtabar, D. Görür, R. Pascanu, and H. Ghasemzadeh · 2021
Closest in time.
Discovering diverse solutions in deep reinforcement learning, 2021
T. Osa, V. Tangkaratt, and M. Sugiyama · 2021
Closest in time.
Learning neural network subspaces
M. Wortsman, M. Horton, C. Guestrin, A. Farhadi, and M. Rastegari · 2021
Closest in time.
Robust reinforcement learning on state observations with learned optimal adversary
H. Zhang, H. Chen, D. S. Boning, and C. Hsieh · 2021
Closest in time.