Fetching the paper…
Reading the bibliography…
A long-standing goal of reinforcement learning is to acquire agents that can learn on training tasks and generalize well on unseen tasks that may share a similar dynamic but with different reward functions.
Mnih V, Badia A P, Mirza M, et al. Asynchronous methods for deep reinforcement learning. In: Proceedings of the 33rd International Conference on Machine Learning, New York, 2016. 1928–1937
1937
Earlier work this paper cites.
Kakade S, Langford J. Approximately optimal approximate reinforcement learning. In: Proceedings of the 19th International Conference on Machine Learning, Sydney, 2002. 267–274
2002
Earlier work this paper cites.
Van der Maaten L, Hinton G. Visualizing data using t-sne. J Mach Learn Res, 2008, 9: 2579–2605
2008
Earlier work this paper cites.
Todorov E, Erez T, Tassa Y. Mujoco: a physics engine for model-based control. In: Proceedings of the 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, Vilamoura, 2012. 5026–5033
2012
Earlier work this paper cites.
Todorov E, Erez T, Tassa Y. Mujoco: a physics engine for model-based control. In: Proceedings of the 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, Vilamoura, 2012. 5026–5033
2012
Earlier work this paper cites.
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
Silver D, Huang A, Maddison C J, et al. Mastering the game of Go with deep neural networks and tree search. Nature, 2016, 529(7587): 484–489
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Finn C, Abbeel P, Levine S. Model-agnostic meta-learning for fast adaptation of deep networks. In: Proceedings of the 34th International Conference on Machine Learning, Sydney, 2017. 1126–1135
2017
Earlier work this paper cites.
Ha D, Schmidhuber J. World models. 2018. ArXiv:1803.10122
2018
Earlier work this paper cites.
Tassa Y, Doron Y, Muldal A, et al. Deepmind control suite. 2018. ArXiv:1801.00690
2018
Earlier work this paper cites.
Nagabandi A, Clavera I, Liu S, et al. Learning to adapt in dynamic, real-world environments through meta-reinforcement learning. In: Proceedings of the 6th International Conference on Learning Representations, Vancouver, 2018. 1–17
2018
Earlier work this paper cites.
Sutton R S, Barto A G. Reinforcement Learning: An Introduction. Cambridge: MIT Press, 2018. 1–622
2018
Earlier work this paper cites.
Hafner D, Lillicrap T, Fischer I, et al. Learning latent dynamics for planning from pixels. In: Proceedings of the 36th International Conference on Machine Learning, Long Beach, 2019. 2555–2565
2019
Earlier work this paper cites.
Cobbe K, Klimov O, Hesse C, et al. Quantifying generalization in reinforcement learning. In: Proceedings of the 36th International Conference on Machine Learning, Long Beach, 2019. 1282–1289
2019
Earlier work this paper cites.
Lee K, Lee K, Shin J, et al. Network randomization: a simple technique for generalization in deep reinforcement learning. In: Proceedings of the 7th International Conference on Learning Representations, New Orleans, 2019. 1–12
2019
Earlier work this paper cites.
Rakelly K, Zhou A, Finn C, et al. Efficient off-policy meta-reinforcement learning via probabilistic context variables. In: Proceedings of the 36th International Conference on Machine Learning, Long Beach, 2019. 5331–5340
2019
Earlier work this paper cites.
Hafner D, Lillicrap T, Ba J, et al. Dream to control: learning behaviors by latent imagination. In: Proceedings of the 8th International Conference on Learning Representations, Addis Ababa, 2020. 1–15
2020
Cited alongside, same era.
Zintgraf L, Shiarlis K, Igl M, et al. Varibad: a very good method for bayes-adaptive deep rl via meta-learning. In: Proceedings of the 8th International Conference on Learning Representations, Addis Ababa, 2020. 1–20
2020
Cited alongside, same era.
Song X, Jiang Y, Tu S, et al. Observational overfitting in reinforcement learning. In: Proceedings of the 8th International Conference on Learning Representations, Addis Ababa, 2020. 1–12
2020
Cited alongside, same era.
Wang K, Kang B, Shao J, et al. Improving generalization in reinforcement learning with mixture regularization. In: Proceedings of the 34th Advances in Neural Information Processing Systems, 2020. 7968–7978
2020
Cited alongside, same era.
Fu X, Yang G, Agrawal P, et al. Learning task informed abstractions. In: Proceedings of the 38th International Conference on Machine Learning, 2021. 3480–3491
2021
Later among the works it cites.
Yarats D, Zhang A, Kostrikov I, et al. Improving sample efficiency in model-free reinforcement learning from images. In: Proceedings of the 35th AAAI Conference on Artificial Intelligence, 2021. 10674–10681
2021
Later among the works it cites.
2021
Later among the works it cites.
Luo F M, Jiang S, Yu Y, et al. Adapt to environment sudden changes by learning a context sensitive policy. In: Proceedings of the 36th AAAI Conference on Artificial Intelligence, Vancouver, 2022. 7637–7646
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lee K, Seo Y, Lee S, et al. Context-aware dynamics model for generalization in model-based reinforcement learning. In: Proceedings of the 37th International Conference on Machine Learning, 2020. 5757–5766
2020
Cited alongside, same era.
Yang R, Xu H, Wu Y, et al. Multi-task reinforcement learning with soft modularization. In: Proceedings of the 34th Advances in Neural Information Processing Systems, 2020. 4767–4777
2020
Cited alongside, same era.
Zintgraf L, Shiarlis K, Igl M, et al. Varibad: a very good method for bayes-adaptive deep rl via meta-learning. In: Proceedings of the 8th International Conference on Learning Representations, Addis Ababa, 2020. 1–20
2020
Cited alongside, same era.
Sekar R, Rybkin O, Daniilidis K, et al. Planning to explore via self-supervised world models. In: Proceedings of the 37th International Conference on Machine Learning, 2020. 8583–8592
2020
Cited alongside, same era.
Khosla P, Teterwak P, Wang C, et al. Supervised contrastive learning. In: Proceedings of the 34th Advances in Neural Information Processing Systems, 2020. 18661–18673
2020
Cited alongside, same era.
Laskin M, Srinivas A, Abbeel P. Curl: contrastive unsupervised representations for reinforcement learning. In: Proceedings of the 37th International Conference on Machine Learning, 2020. 5639–5650
2020
Cited alongside, same era.
Touati A, Ollivier Y. Learning one representation to optimize all rewards. In: Proceedings of the 35th Advances in Neural Information Processing Systems, 2021. 13–23
2021
Cited alongside, same era.
Raileanu R, Fergus R. Decoupling value and policy for generalization in reinforcement learning. In: Proceedings of the 38th International Conference on Machine Learning, 2021. 8787–8798
2021
Cited alongside, same era.
Lee K H, Nachum O, Yang M S, et al. Multi-game decision transformers. In: Proceedings of the 36th Advances in Neural Information Processing Systems, New Orleans, 2022. 27921–27936
2022
Later among the works it cites.
2022
Later among the works it cites.
Deng F, Jang I, Ahn S. Dreamerpro: reconstruction-free model-based reinforcement learning with prototypical representations. In: Proceedings of the 39th International Conference on Machine Learning, Baltimore, 2022. 4956–4975
2022
Later among the works it cites.
Wang T, Du S, Torralba A, et al. Denoised mdps: learning world models better than the world itself. In: Proceedings of the 39th International Conference on Machine Learning, Baltimore, 2022. 22591–22612
2022
Later among the works it cites.
Seo Y, Lee K, James S L, et al. Reinforcement learning with action-free pre-training from videos. In: Proceedings of the 39th International Conference on Machine Learning, Baltimore, 2022. 19561–19579
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Touati A, Rapin J, Ollivier Y. Does zero-shot reinforcement learning exist? In: Proceedings of the 11th International Conference on Learning Representations, Kigali, 2023. 1–18
2023
Closest in time.
Xue Z, Cai Q, Liu S, et al. State regularized policy optimization on data with dynamics shift. In: Proceedings of the 37th Advances in Neural Information Processing Systems, New Orleans, 2023. 32926–32937
2023
Closest in time.
2023
Closest in time.
Young K J, Ramesh A, Kirsch L, et al. The benefits of model-based generalization in reinforcement learning. In: Proceedings of the 40th International Conference on Machine Learning, Honolulu, 2023. 40254–40276
2023
Closest in time.
Zheng R, Wang X, Sun Y, et al. Taco: temporal latent action-driven contrastive loss for visual reinforcement learning. In: Proceedings of the 37th Advances in Neural Information Processing Systems, 2023. 48203–48225
2023
Closest in time.
Rimon Z, Jurgenson T, Krupnik O, et al. Mamba: an effective world model approach for meta-reinforcement learning. In: Proceedings of the 12th International Conference on Learning Representations, Vienna, 2024. 1–25
2024
Closest in time.