Fetching the paper…
Reading the bibliography…
A zero-shot RL agent is an agent that can solve any RL task in a given environment, instantly with no additional planning or learning, after an initial reward-free learning phase.
Improving generalization for temporal difference learning: The successor representation
Peter Dayan · 1993
Earlier work this paper cites.
Stochastic differential equations
Bernt Øksendal · 1998
Earlier work this paper cites.
Structure in the space of value functions
David Foster and Peter Dayan · 2002
Earlier work this paper cites.
On spectral graph drawing
Yehuda Koren · 2003
Earlier work this paper cites.
Proto-value functions: A Laplacian framework for learning representation and control in Markov decision processes
Sridhar Mahadevan and Mauro Maggioni · 2007
Earlier work this paper cites.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Markov chains and mixing times
David A Levin, Yuval Peres, and Elisabeth L Wilmer · 2009
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Richard S Sutton, Joseph Modayil, Michael Delp, Thomas Degris, Patrick M Pilarski, Adam White, and Doina Precup · 2011
Earlier work this paper cites.
Learning parameterized skills
Bruno Castro da Silva, George Konidaris, and Andrew G Barto · 2012
Earlier work this paper cites.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2016
Earlier work this paper cites.
Hindsight experience replay
Marcin Andrychowicz, Dwight Crow, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel, and Wojciech Zaremba · 2017
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
André Barreto, Will Dabney, Rémi Munos, Jonathan J Hunt, Tom Schaul, David Silver, and Hado P van Hasselt · 2017
Earlier work this paper cites.
A Laplacian framework for option discovery in reinforcement learning
Marlos C Machado, Marc G Bellemare, and Michael Bowling · 2017
Earlier work this paper cites.
Zero-shot task generalization with multi-task deep reinforcement learning
Junhyuk Oh, Satinder Singh, Honglak Lee, and Pushmeet Kohli · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Cited alongside, same era.
Deep reinforcement learning with successor features for navigation across similar environments
Jingwei Zhang, Jost Tobias Springenberg, Joschka Boedecker, and Wolfram Burgard · 2017
Cited alongside, same era.
Universal successor features approximators
Diana Borsa, André Barreto, John Quan, Daniel Mankowitz, Rémi Munos, Hado van Hasselt, David Silver, and Tom Schaul · 2018
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine · 2018
Zero-shot reinforcement learning with deep attention convolutional neural networks
Sahika Genc, Sunil Mallya, Sravan Bodapati, Tao Sun, and Yunzhe Tao · 2020
Later among the works it cites.
Bootstrap your own latent-a new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al · 2020
Later among the works it cites.
Universal successor features for transfer reinforcement learning
Chen Ma, Dylan R Ashley, Junfeng Wen, and Yoshua Bengio · 2020
Later among the works it cites.
Model-based reinforcement learning: A survey
Thomas M Moerland, Joost Broekens, and Catholijn M Jonker · 2020
Later among the works it cites.
Learning successor states and goal-dependent values: A mathematical viewpoint
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Hierarchical reinforcement learning for zero-shot generalization with subtask dependencies
Sungryull Sohn, Junhyuk Oh, and Honglak Lee · 2018
Cited alongside, same era.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al · 2018
Cited alongside, same era.
The Laplacian in RL: Learning representations with efficient approximations
Yifan Wu, George Tucker, and Ofir Nachum · 2018
Cited alongside, same era.
Disentangled cumulants help successor representations transfer to new tasks
Christopher Grimm, Irina Higgins, Andre Barreto, Denis Teplyashin, Markus Wulfmeier, Tim Hertweck, Raia Hadsell, and Satinder Singh · 2019
Cited alongside, same era.
Léonard Blier, Corentin Tallec, and Yann Ollivier · 2021
Later among the works it cites.
APS: Active pretraining with successor features
Hao Liu and Pieter Abbeel · 2021
Later among the works it cites.
Urlb: Unsupervised reinforcement learning benchmark
Michael Laskin, Denis Yarats, Hao Liu, Kimin Lee, Albert Zhan, Kevin Lu, Catherine Cang, Lerrel Pinto, and Pieter Abbeel · 2021
Later among the works it cites.
Learning one representation to optimize all rewards
Ahmed Touati and Yann Ollivier · 2021
Later among the works it cites.
Reinforcement learning with prototypical representations
Denis Yarats, Rob Fergus, Alessandro Lazaric, and Lerrel Pinto · 2021
Later among the works it cites.
Randall Balestriero and Yann LeCun · 2022
Closest in time.
Cic: Contrastive intrinsic control for unsupervised skill discovery
Michael Laskin, Hao Liu, Xue Bin Peng, Denis Yarats, Aravind Rajeswaran, and Pieter Abbeel · 2022
Closest in time.
Understanding and preventing capacity loss in reinforcement learning
Clare Lyle, Mark Rowland, and Will Dabney · 2022
Closest in time.
Spectral decomposition representation for reinforcement learning
Tongzheng Ren, Tianjun Zhang, Lisa Lee, Joseph E Gonzalez, Dale Schuurmans, and Bo Dai · 2022
Closest in time.
Deep contrastive learning is provably (almost) principal component analysis
Yuandong Tian · 2022
Closest in time.