Fetching the paper…
Reading the bibliography…
Zero-shot generalization (ZSG) to unseen dynamics is a major challenge for creating generally capable embodied agents.
Neuronlike adaptive elements that can solve difficult learning control problems
A. G. Barto, R. S. Sutton, and C. W. Anderson · 1983
Earlier work this paper cites.
Neuroevolutionary reinforcement learning for generalized helicopter control
R. Koppejan and S. Whiteson · 2009
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
D. Kingma and M. Welling · 2014
Earlier work this paper cites.
Contextual markov decision processes
A. Hallak, D. Di Castro, and S. Mannor · 2015
Earlier work this paper cites.
Charlie Beattie, Joel Z. Leibo, Denis Teplyashin, Tom Ward, Marcus Wainwright, Heinrich Küttler, Andrew Lefrancq, Simon Green, Víctor Valdés, Amir Sadik, Julian Schrittwieser, Keith Anderson, Sarah York, Max Cant, Adam Cain, Adrian Bolton, Stephen Gaffney, Helen King, Demis Hassabis, Shane Legg, and Stig Petersen · 2016
Earlier work this paper cites.
RL$ˆ2$: Fast reinforcement learning via slow reinforcement learning
Y. Duan, J. Schulman, X. Chen, P. Bartlett, I. Sutskever, and P. Abbeel · 2016
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel · 2017
Earlier work this paper cites.
Learning to reinforcement learn
J. Wang, Z. Kurth-Nelson, H. Soyer, J. Leibo, D. Tirumala, R. Munos, C. Blundell, D. Kumaran, and M. Botvinick · 2017
Earlier work this paper cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
K. Chua, R. Calandra, R. McAllister, and S. Levine · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Markov decision processes with continuous side information
A. Modi, N. Jiang, S. Singh, and A. Tewari · 2018
Earlier work this paper cites.
Sim-to-real transfer of robotic control with dynamics randomization
X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel · 2018
Earlier work this paper cites.
Deepmind control suite
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. L. Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, T. Lillicrap, and M. Riedmiller · 2018
Earlier work this paper cites.
Provably efficient RL with rich observations via latent state decoding
S. Du, A. Krishnamurthy, N. Jiang, A. Agarwal, M. Dudik, and J. Langford · 2019
Earlier work this paper cites.
The minerl 2019 competition on sample efficient reinforcement learning using human priors
William H Guss, Cayden Codel, Katja Hofmann, Brandon Houghton, Noboru Kuno, Stephanie Milani, Sharada Mohanty, Diego Perez Liebana, Ruslan Salakhutdinov, Nicholay Topin, et al · 2019
Cited alongside, same era.
Learning latent dynamics for planning from pixels
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson · 2019
Cited alongside, same era.
Learning to adapt in dynamic, real-world environments through meta-reinforcement learning
A. Nagabandi, I. Clavera, S. Liu, R. S. Fearing, P. Abbeel, S. Levine, and C. Finn · 2019
Cited alongside, same era.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
K. Rakelly, A. Zhou, C. Finn, S. Levine, and D. Quillen · 2019
Cited alongside, same era.
Learning dexterous in-hand manipulation
M. Andrychowicz, B. Baker, M. Chociej, R. Józefowicz, B. McGrew, J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, J. Schneider, S. Sidor, J. Tobin, P. Welinder, L. Weng, and W. Zaremba · 2020
Learning domain-independent policies for open list selection
A. Biedenkapp, D. Speck, S. Sievers, F. Hutter, M. Lindauer, and J. Seipp · 2022
Later among the works it cites.
A relational intervention approach for unsupervised dynamics generalization in model-based reinforcement learning
J. Guo, M. Gong, and D. Tao · 2022
Later among the works it cites.
Transformers are meta-reinforcement learners
L. C. Melo · 2022
Later among the works it cites.
Automated reinforcement learning (AutoRL): A survey and open problems
J. Parker-Holder, R. Rajan, X. Song, A. Biedenkapp, Y. Miao, T. Eimer, B. Zhang, V. Nguyen, R. Calandra, A. Faust, F. Hutter, and M. Lindauer · 2022
Later among the works it cites.
Block contextual mdps for continual learning
S. Sodhani, F. Meier, J. Pineau, and A. Zhang · 2022
Later among the works it cites.
Daydreamer: World models for physical robot learning
P. Wu, A. Escontrela, D. Hafner, P. Abbeel, and K. Goldberg · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Zero-shot terrain generalization for visual locomotion policies
A. Escontrela, G. Yu, P. Xu, A. Iscen, and J. Tan · 2020
Cited alongside, same era.
Dream to control: Learning behaviors by latent imagination
D. Hafner, T. P. Lillicrap, J. Ba, and M. Norouzi · 2020
Cited alongside, same era.
Context-aware dynamics model for generalization in model-based reinforcement learning
K. Lee, Y. Seo, S. Lee, H. Lee, and J. Shin · 2020
Cited alongside, same era.
Generalized hidden parameter mdps: Transferable model-based rl in a handful of trials
C. Perez, F. P. Such, and T. Karaletsos · 2020
Cited alongside, same era.
Deep reinforcement learning at the edge of the statistical precipice
R. Agarwal, M. Schwarzer, P. Samuel Castro, A. C. Courville, and M. G. Bellemare · 2021
Cited alongside, same era.
Procedural generalization by planning with self-supervised world models
Ankesh Anand, Jacob C Walker, Yazhe Li, Eszter Vértes, Julian Schrittwieser, Sherjil Ozair, Theophane Weber, and Jessica B Hamrick · 2021
Cited alongside, same era.
Augmented world models facilitate zero-shot dynamics generalization from a single offline environment
P. J. Ball, C. Lu, J. Parker-Holder, and S. Roberts · 2021
Cited alongside, same era.
Later among the works it cites.
A survey of meta-reinforcement learning
J. Beck, R. Vuorio, E. Z. Liu, Z. Xiong, L. Zintgraf, C. Finn, and S. Whiteson · 2023
Later among the works it cites.
Contextualize me – the case for context in reinforcement learning
C. Benjamins, T. Eimer, F. Schubert, A. Mohan, S. Döhler, A. Biedenkapp, B. Rosenhan, F. Hutter, and M. Lindauer · 2023
Later among the works it cites.
Dynamics generalisation in reinforcement learning via adaptive context-aware policies
M. Beukman, D. Jarvis, R. Klein, S. James, and B. Rosman · 2023
Later among the works it cites.
Mastering diverse domains through world models
D. Hafner, J. Pasukonis, J. Ba, and T. P. Lillicrap · 2023
Later among the works it cites.
A survey of zero-shot generalisation in deep reinforcement learning
R. Kirk, A. Zhang, E. Grefenstette, and T. Rocktäschel · 2023
Later among the works it cites.
Dream to adapt: Meta reinforcement learning by latent context imagination and MDP imagination
L. Wen, S. Zhang, H. E. Tseng, and H. Peng · 2023
Later among the works it cites.
The benefits of model-based generalization in reinforcement learning
K. J. Young, A. Ramesh, L. Kirsch, and J. Schmidhuber · 2023
Later among the works it cites.
TD-MPC2: Scalable, robust world models for continuous control
N. Hansen, H. Su, and X. Wang · 2024
Closest in time.