Fetching the paper…
Reading the bibliography…
It is of significance for an agent to learn a widely applicable and general-purpose policy that can achieve diverse goals including images and text descriptions.
Self-supervised learning of image embedding for continuous control
Florensa, C.; Degrave, J.; Heess, N.; Springenberg, J. T.; and Riedmiller, M. 2019 · 1901
Earlier work this paper cites.
Skew-fit: State-covering self-supervised reinforcement learning
Pong, V. H.; Dalal, M.; Lin, S.; Nair, A.; Bahl, S.; and Levine, S. 2019 · 1903
Earlier work this paper cites.
Efficient Exploration via State Marginal Matching
Lee, L.; Eysenbach, B.; Parisotto, E.; Xing, E. P.; Levine, S.; and Salakhutdinov, R. 2019 · 1906
Earlier work this paper cites.
MediaPipe: A Framework for Building Perception Pipelines
Lugaresi, C.; Tang, J.; Nash, H.; McClanahan, C.; Uboweja, E.; Hays, M.; Zhang, F.; Chang, C.; Yong, M. G.; Lee, J.; Chang, W.; Hua, W.; Georg, M.; and Grundmann, M. 2019 · 1906
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
Berner, C.; Brockman, G.; Chan, B.; Cheung, V.; Debiak, P.; Dennison, C.; Farhi, D.; Fischer, Q.; Hashme, S.; Hesse, C.; et al. 2019 · 1912
Earlier work this paper cites.
Explore, Discover and Learn: Unsupervised Discovery of State-Covering Skills
Campos, V. A.; Trott, A.; Xiong, C.; Socher, R.; i Nieto, X. G.; and Torres, J. 2020 · 2002
Earlier work this paper cites.
Agent57: Outperforming the atari human benchmark
Badia, A. P.; Piot, B.; Kapturowski, S.; Sprechmann, P.; Vitvitskyi, A.; Guo, D.; and Blundell, C. 2020 · 2003
Earlier work this paper cites.
The IM algorithm: a variational approach to information maximization
Barber, D.; and Agakov, F. V. 2003 · 2003
Earlier work this paper cites.
Knowledge Distillation Meets Self-Supervision
Xu, G.; Liu, Z.; Li, X.; and Loy, C. C. 2020 · 2006
Earlier work this paper cites.
Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning
Pitis, S.; Chan, H.; Zhao, S.; Stadie, B. C.; and Ba, J. 2020 · 2007
Earlier work this paper cites.
Grimgep: learning progress for robust goal sampling in visual deep reinforcement learning
Kovač, G.; Laversanne-Finot, A.; and Oudeyer, P.-Y. 2020 · 2008
Earlier work this paper cites.
Dynamics Generalization via Information Bottleneck in Deep Reinforcement Learning
Lu, X.; Lee, K.; Abbeel, P.; and Tiomkin, S. 2020 · 2008
Earlier work this paper cites.
Action and Perception as Divergence Minimization
Hafner, D.; Ortega, P. A.; Ba, J.; Parr, T.; Friston, K. J.; and Heess, N. 2020 · 2009
Earlier work this paper cites.
An intuitive proof of the data processing inequality
Beaudry, N. J.; and Renner, R. 2011 · 2011
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
Todorov, E.; Erez, T.; and Tassa, Y. 2012 · 2012
Earlier work this paper cites.
Empowerment–an introduction
Salge, C.; Glackin, C.; and Polani, D. 2014 · 2014
Earlier work this paper cites.
Universal value function approximators
Schaul, T.; Horgan, D.; Gregor, K.; and Silver, D. 2015 · 2015
Earlier work this paper cites.
Concrete problems in AI safety
Amodei, D.; Olah, C.; Steinhardt, J.; Christiano, P. F.; Schulman, J.; and Mané, D. 2016 · 2016
Cited alongside, same era.
OpenAI Gym
Brockman, G.; Cheung, V.; Pettersson, L.; Schneider, J.; Schulman, J.; Tang, J.; and Zaremba, W. 2016 · 2016
Cited alongside, same era.
Hindsight experience replay
Andrychowicz, M.; Wolski, F.; Ray, A.; Schneider, J.; Fong, R.; Welinder, P.; McGrew, B.; Tobin, J.; Abbeel, O. P.; and Zaremba, W. 2017 · 2017
Cited alongside, same era.
Variational Intrinsic Control
Gregor, K.; Rezende, D. J.; and Wierstra, D. 2017 · 2017
Cited alongside, same era.
Darla: Improving zero-shot transfer in reinforcement learning
Higgins, I.; Pal, A.; Rusu, A.; Matthey, L.; Burgess, C.; Pritzel, A.; Botvinick, M.; Blundell, C.; and Lerchner, A. 2017 · 2017
Cited alongside, same era.
Visual reinforcement learning with imagined goals
Nair, A. V.; Pong, V.; Dalal, M.; Bahl, S.; Lin, S.; and Levine, S. 2018 · 2018
Later among the works it cites.
Unsupervised learning of goal spaces for intrinsically motivated goal exploration
Péré, A.; Forestier, S.; Sigaud, O.; and Oudeyer, P.-Y. 2018 · 2018
Later among the works it cites.
Temporal difference models: Model-free deep rl for model-based control
Pong, V.; Gu, S.; Dalal, M.; and Levine, S. 2018 · 2018
Later among the works it cites.
Time-contrastive networks: Self-supervised learning from video
Sermanet, P.; Lynch, C.; Chebotar, Y.; Hsu, J.; Jang, E.; Schaal, S.; Levine, S.; and Brain, G. 2018 · 2018
Later among the works it cites.
Learning goal embeddings via self-play for hierarchical reinforcement learning
Sukhbaatar, S.; Denton, E.; Szlam, A.; and Fergus, R. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Levy, A.; Konidaris, G.; Platt, R.; and Saenko, K. 2017 · 2017
Cited alongside, same era.
Data-efficient deep reinforcement learning for dexterous manipulation
Popov, I.; Heess, N.; Lillicrap, T.; Hafner, R.; Barth-Maron, G.; Vecerik, M.; Lampe, T.; Tassa, Y.; Erez, T.; and Riedmiller, M. 2017 · 2017
Cited alongside, same era.
Proximal Policy Optimization Algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Cited alongside, same era.
Intrinsic motivation and automatic curricula via asymmetric self-play
Sukhbaatar, S.; Lin, Z.; Kostrikov, I.; Synnaeve, G.; Szlam, A.; and Fergus, R. 2017 · 2017
Cited alongside, same era.
Domain randomization for transferring deep neural networks from simulation to the real world
Tobin, J.; Fong, R.; Ray, A.; Schneider, J.; Zaremba, W.; and Abbeel, P. 2017 · 2017
Cited alongside, same era.
Variational Option Discovery Algorithms
Achiam, J.; Edwards, H.; Amodei, D.; and Abbeel, P. 2018 · 2018
Cited alongside, same era.
Gep-pg: Decoupling exploration and exploitation in deep reinforcement learning algorithms
Colas, C.; Sigaud, O.; and Oudeyer, P.-Y. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Dynamical Distance Learning for Semi-Supervised and Unsupervised Skill Discovery
Hartikainen, K.; Geng, X.; Haarnoja, T.; and Levine, S. 2019 · 2019
Later among the works it cites.
Unsupervised curricula for visual meta-reinforcement learning
Jabri, A.; Hsu, K.; Gupta, A.; Eysenbach, B.; Levine, S.; and Finn, C. 2019 · 2019
Later among the works it cites.
Plan Online, Learn Offline: Efficient Learning and Exploration via Model-Based Control
Lowrey, K.; Rajeswaran, A.; Kakade, S. M.; Todorov, E.; and Mordatch, I. 2019 · 2019
Later among the works it cites.
Contextual Imagined Goals for Self-Supervised Robotic Learning
Nair, A.; Bahl, S.; Khazatsky, A.; Pong, V.; Berseth, G.; and Levine, S. 2019 · 2019
Later among the works it cites.
A practical approach to insertion with variable socket position using deep reinforcement learning
Vecerik, M.; Sushkov, O.; Barker, D.; Rothörl, T.; Hester, T.; and Scholz, J. 2019 · 2019
Later among the works it cites.
Unsupervised Control Through Non-Parametric Discriminative Rewards
Warde-Farley, D.; de Wiele, T. V.; Kulkarni, T. D.; Ionescu, C.; Hansen, S.; and Mnih, V. 2019 · 2019
Later among the works it cites.
Language as a Cognitive Tool to Imagine Goals in Curiosity Driven Exploration
Colas, C.; Karch, T.; Lair, N.; Dussoux, J.; Moulin-Frier, C.; Dominey, P. F.; and Oudeyer, P. 2020 · 2020
Later among the works it cites.
CURL: Contrastive Unsupervised Representations for Reinforcement Learning
Laskin, M.; Srinivas, A.; and Abbeel, P. 2020 · 2020
Later among the works it cites.
Dynamics-Aware Unsupervised Discovery of Skills
Sharma, A.; Gu, S.; Levine, S.; Kumar, V.; and Hausman, K. 2020 · 2020
Later among the works it cites.
Learning transitional skills with intrinsic motivation
Tian, Q.; Liu, J.; and Wang, D. 2020 · 2020
Later among the works it cites.
Return-Based Contrastive Representation Learning for Reinforcement Learning
Liu, G.; Zhang, C.; Zhao, L.; Qin, T.; Zhu, J.; Li, J.; Yu, N.; and Liu, T. 2021 · 2021
Closest in time.
Hierarchical Reinforcement Learning By Discovering Intrinsic Options
Zhang, J.; Yu, H.; and Xu, W. 2021 · 2021
Closest in time.