Fetching the paper…
Reading the bibliography…
How can artificial agents learn to solve many diverse tasks in complex visual environments in the absence of any supervision? We decompose this question into two problems: discovering new goals and learning to reliably achieve them.
On a measure of the information provided by an experiment
D. V. Lindley et al · 1956
Earlier work this paper cites.
Curious model-building control systems
J. Schmidhuber · 1991
Earlier work this paper cites.
Learning to achieve goals
L. P. Kaelbling · 1993
Earlier work this paper cites.
Planning to be surprised: Optimal bayesian exploration in dynamic environments
Y. Sun, F. Gomez, and J. Schmidhuber · 2011
Earlier work this paper cites.
Finding new facts; thinking new thoughts
L. Schulz · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
D. J. Rezende, S. Mohamed, and D. Wierstra · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
Universal value function approximators
T. Schaul, D. Horgan, K. Gregor, and D. Silver · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
M. Watter, J. T. Springenberg, J. Boedecker, and M. Riedmiller · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Earlier work this paper cites.
K. Gregor, D. J. Rezende, and D. Wierstra · 2016
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
B. Lakshminarayanan, A. Pritzel, and C. Blundell · 2016
Earlier work this paper cites.
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, P. Abbeel, and W. Zaremba · 2017
Earlier work this paper cites.
Deep visual foresight for planning robot motion
C. Finn and S. Levine · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Earlier work this paper cites.
Variational option discovery algorithms
J. Achiam, H. Edwards, D. Amodei, and P. Abbeel · 2018
Earlier work this paper cites.
Learning and querying fast generative models for reinforcement learning
L. Buesing, T. Weber, S. Racaniere, S. Eslami, D. Rezende, D. P. Reichert, F. Viola, F. Besse, K. Gregor, D. Hassabis, et al · 2018
Cited alongside, same era.
Exploration by random network distillation
Y. Burda, H. Edwards, A. Storkey, and O. Klimov · 2018
Cited alongside, same era.
Visual foresight: Model-based deep reinforcement learning for vision-based robotic control
F. Ebert, C. Finn, S. Dasari, A. Xie, A. Lee, and S. Levine · 2018
Cited alongside, same era.
Diversity is all you need: learning skills without a reward function
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine · 2018
Cited alongside, same era.
Automatic goal generation for reinforcement learning agents
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
A. Gupta, V. Kumar, C. Lynch, S. Levine, and K. Hausman · 2019
Later among the works it cites.
Dream to control: Learning behaviors by latent imagination
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi · 2019
Later among the works it cites.
Self-supervised exploration via disagreement
D. Pathak, D. Gandhi, and A. Gupta · 2019
Later among the works it cites.
Skew-fit: State-covering self-supervised reinforcement learning
V. H. Pong, M. Dalal, S. Lin, A. Nair, S. Bahl, and S. Levine · 2019
Later among the works it cites.
Ready policy one: World building through active learning
P. Ball, J. Parker-Holder, A. Pacchiano, K. Choromanski, and S. Roberts · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Florensa, D. Held, X. Geng, and P. Abbeel · 2018
Cited alongside, same era.
Learning latent dynamics for planning from pixels
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson · 2018
Cited alongside, same era.
Visual reinforcement learning with imagined goals
A. V. Nair, V. Pong, M. Dalal, S. Bahl, S. Lin, and S. Levine · 2018
Cited alongside, same era.
Zero-shot visual imitation
D. Pathak, P. Mahmoudieh, G. Luo, P. Agrawal, D. Chen, Y. Shentu, E. Shelhamer, J. Malik, A. A. Efros, and T. Darrell · 2018
Cited alongside, same era.
Model-based active exploration
P. Shyam, W. Jaśkowski, and F. Gomez · 2018
Cited alongside, same era.
Intrinsic motivation and automatic curricula via asymmetric self-play
S. Sukhbaatar, Z. Lin, I. Kostrikov, G. Synnaeve, A. Szlam, and R. Fergus · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Cited alongside, same era.
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. d. L. Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, et al · 2018
Cited alongside, same era.
B. Bucher, K. Schmeckpeper, N. Matni, and K. Daniilidis · 2020
Later among the works it cites.
Explore, discover and learn: Unsupervised discovery of state-covering skills
V. Campos, A. Trott, C. Xiong, R. Socher, X. Giró-i Nieto, and J. Torres · 2020
Later among the works it cites.
Variational empowerment as representation learning for goal-based reinforcement learning
J. Choi, A. Sharma, S. Levine, H. Lee, and S. S. Gu · 2020
Later among the works it cites.
Mastering atari with discrete world models
D. Hafner, T. Lillicrap, M. Norouzi, and J. Ba · 2020
Later among the works it cites.
Dynamical distance learning for semi-supervised and unsupervised skill discovery
K. Hartikainen, X. Geng, T. Haarnoja, and S. Levine · 2020
Later among the works it cites.
Planning to explore via self-supervised world models
R. Sekar, O. Rybkin, K. Daniilidis, P. Abbeel, D. Hafner, and D. Pathak · 2020
Later among the works it cites.
Model-based visual planning with self-supervised functional distances
S. Tian, S. Nair, F. Ebert, S. Dasari, B. Eysenbach, C. Finn, and S. Levine · 2020
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine · 2020
Later among the works it cites.
Automatic curriculum learning through value disagreement
Y. Zhang, P. Abbeel, and L. Pinto · 2020
Later among the works it cites.
Actionable models: Unsupervised offline reinforcement learning of robotic skills
Y. Chebotar, K. Hausman, Y. Lu, T. Xiao, D. Kalashnikov, J. Varley, A. Irpan, B. Eysenbach, R. Julian, C. Finn, et al · 2021
Closest in time.
Asymmetric self-play for automatic goal discovery in robotic manipulation
O. OpenAI, M. Plappert, R. Sampedro, T. Xu, I. Akkaya, V. Kosaraju, P. Welinder, R. D’Sa, A. Petron, H. P. d. O. Pinto, et al · 2021
Closest in time.
Clockwork variational autoencoders
V. Saxena, J. Ba, and D. Hafner · 2021
Closest in time.
Mastering visual continuous control: Improved data-augmented reinforcement learning
D. Yarats, R. Fergus, A. Lazaric, and L. Pinto · 2021
Closest in time.