Fetching the paper…
Reading the bibliography…
There has been a recent surge of interest in developing generally-capable agents that can adapt to new tasks without additional training in the environment.
The theory of statistical decision
Leonard J Savage · 1951
Earlier work this paper cites.
Intrinsic motivation and self-determination in human behavior
Edward L Deci and Richard M Ryan · 1985
Earlier work this paper cites.
Reinforcement learning in Markovian and non-Markovian environments
Jürgen Schmidhuber · 1990
Earlier work this paper cites.
A possibility for implementing curiosity and boredom in model-building neural controllers
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Richard S Sutton · 1991
Earlier work this paper cites.
Purposive behavior acquisition for a real robot by vision-based reinforcement learning
Minoru Asada, Shoichi Noda, Sukoya Tawaratsumida, and Koh Hosoda · 1996
Earlier work this paper cites.
Active learning with statistical models
David A Cohn, Zoubin Ghahramani, and Michael I Jordan · 1996
Earlier work this paper cites.
Evolutionary robotics and the radical envelope-of-noise hypothesis
Nick Jakobi · 1997
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L Littman, and Anthony R Cassandra · 1998
Earlier work this paper cites.
Optimization of conditional value-at-risk
R Tyrrell Rockafellar, Stanislav Uryasev, et al · 2000
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Improving noise
Ken Perlin · 2002
Earlier work this paper cites.
Empowerment: A universal agent-centric measure of control
Alexander S Klyubin, Daniel Polani, and Chrystopher L Nehaniv · 2005
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Jürgen Schmidhuber · 2010
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Reverse curriculum generation for reinforcement learning
Carlos Florensa, David Held, Markus Wulfmeier, Michael Zhang, and Pieter Abbeel · 2017
Earlier work this paper cites.
Automated curriculum learning for neural networks
Alex Graves, Marc G Bellemare, Jacob Menick, Remi Munos, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
Deep variational bayes filters: Unsupervised learning of state space models from raw data
Maximilian Karl, Maximilian Soelch, Justin Bayer, and Patrick Van der Smagt · 2017
Earlier work this paper cites.
Count-based exploration with neural density models
Georg Ostrovski, Marc G Bellemare, Aäron Oord, and Rémi Munos · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Cited alongside, same era.
Domain randomization for transferring deep neural networks from simulation to the real world
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel · 2017
Cited alongside, same era.
Automatic goal generation for reinforcement learning agents
Carlos Florensa, David Held, Xinyang Geng, and Pieter Abbeel · 2018
Cited alongside, same era.
Recurrent world models facilitate policy evolution
David Ha and Jürgen Schmidhuber · 2018
Cited alongside, same era.
Learning to play with intrinsically-motivated, self-aware agents
Nick Haber, Damian Mrowca, Stephanie Wang, Li F Fei-Fei, and Daniel L Yamins · 2018
Cited alongside, same era.
Mastering Atari, Go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Later among the works it cites.
Planning to explore via self-supervised world models
Ramanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel, Danijar Hafner, and Deepak Pathak · 2020
Later among the works it cites.
MOPO: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Y Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C Courville, and Marc Bellemare · 2021
Later among the works it cites.
Self-paced context evaluation for contextual reinforcement learning
Theresa Eimer, André Biedenkapp, Frank Hutter, and Marius Lindauer · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al · 2018
Cited alongside, same era.
Solving rubik’s cube with a robot hand
Ilge Akkaya, Marcin Andrychowicz, Maciek Chociej, Mateusz Litwin, Bob McGrew, Arthur Petron, Alex Paino, Matthias Plappert, Glenn Powell, Raphael Ribas, et al · 2019
Cited alongside, same era.
Learning latent dynamics for planning from pixels
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson · 2019
Cited alongside, same era.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Cited alongside, same era.
Teacher–student curriculum learning
Tambet Matiisen, Avital Oliver, Taco Cohen, and John Schulman · 2019
Cited alongside, same era.
Benchmarking safe exploration in deep reinforcement learning
Alex Ray, Joshua Achiam, and Dario Amodei · 2019
Cited alongside, same era.
Model-based active exploration
Pranav Shyam, Wojciech Jaśkowski, and Faustino Gomez · 2019
Cited alongside, same era.
Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2021
Later among the works it cites.
URLB: Unsupervised reinforcement learning benchmark
Michael Laskin, Denis Yarats, Hao Liu, Kimin Lee, Albert Zhan, Kevin Lu, Catherine Cang, Lerrel Pinto, and Pieter Abbeel · 2021
Later among the works it cites.
Discovering and achieving goals via world models
Russell Mendonca, Oleh Rybkin, Kostas Daniilidis, Danijar Hafner, and Deepak Pathak · 2021
Later among the works it cites.
Automatic curriculum learning for deep rl: A short survey
Rémy Portelas, Cédric Colas, Lilian Weng, Katja Hofmann, and Pierre-yves Oudeyer · 2021
Later among the works it cites.
Minimax regret optimisation for robust planning in uncertain Markov decision processes
Marc Rigter, Bruno Lacerda, and Nick Hawes · 2021
Later among the works it cites.
Open-ended learning leads to generally capable agents
Adam Stooke, Anuj Mahajan, Catarina Barros, Charlie Deck, Jakob Bauer, Jakub Sygnowski, Maja Trebacz, Max Jaderberg, Michael Mathieu, et al · 2021
Later among the works it cites.
Information prioritization through empowerment in visual model-based RL
Homanga Bharadhwaj, Mohammad Babaeizadeh, Dumitru Erhan, and Sergey Levine · 2022
Later among the works it cites.
Do as I can, not as I say: Grounding language in robotic affordances
Anthony Brohan, Yevgen Chebotar, Chelsea Finn, Karol Hausman, Alexander Herzog, Daniel Ho, Julian Ibarz, Alex Irpan, Eric Jang, Ryan Julian, et al · 2022
Later among the works it cites.
Understanding domain randomization for sim-to-real transfer
Xiaoyu Chen, Jiachen Hu, Chi Jin, Lihong Li, and Liwei Wang · 2022
Later among the works it cites.
Revisiting design choices in offline model-based reinforcement learning
Cong Lu, Philip J Ball, Jack Parker-Holder, Michael A Osborne, and Stephen J Roberts · 2022
Later among the works it cites.
Evolving curricula with regret-based environment design
Jack Parker-Holder, Minqi Jiang, Michael Dennis, Mikayel Samvelyan, Jakob Foerster, Edward Grefenstette, and Tim Rocktäschel · 2022
Later among the works it cites.
A generalist agent
Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, et al · 2022
Later among the works it cites.
RAMBO-RL: Robust adversarial model-based offline reinforcement learning
Marc Rigter, Bruno Lacerda, and Nick Hawes · 2022
Later among the works it cites.
Learning general world models in a handful of reward-free deployments
Yingchen Xu, Jack Parker-Holder, Aldo Pacchiano, Philip Ball, Oleh Rybkin, S Roberts, Tim Rocktäschel, and Edward Grefenstette · 2022
Later among the works it cites.
Mastering the unsupervised reinforcement learning benchmark from pixels
Sai Rajeswar, Pietro Mazzaglia, Tim Verbelen, Alexandre Piché, Bart Dhoedt, Aaron Courville, and Alexandre Lacoste · 2023
Closest in time.
Marc Rigter, Bruno Lacerda, and Nick Hawes · 2023
Closest in time.