Fetching the paper…
Reading the bibliography…
Unsupervised reinforcement learning (URL) poses a promising paradigm to learn useful behaviors in a task-agnostic environment without the guidance of extrinsic rewards to facilitate the fast adaptation of various downstream tasks.
Dyna, an integrated architecture for learning, planning, and reacting
Richard S. Sutton · 1991
Earlier work this paper cites.
Optimization of computer simulation models with rare events
Reuven Y Rubinstein · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Control of systems integrating logic, dynamics, and constraints
Alberto Bemporad and Manfred Morari · 1999
Earlier work this paper cites.
Nearest neighbor estimates of entropy
Harshinder Singh, Neeraj Misra, Vladimir Hnizdo, Adam Fedorowicz, and Eugene Demchuk · 2003
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Pierre-Yves Oudeyer, Frdric Kaplan, and Verena V Hafner · 2007
Earlier work this paper cites.
Mastering atari with discrete world models
Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2010
Earlier work this paper cites.
Seeking multi-thresholds for image segmentation with learning automata
Erik Cuevas, Daniel Zaldivar, and Marco Pérez-Cisneros · 2011
Earlier work this paper cites.
Multiple choice learning: Learning to produce multiple structured outputs
Abner Guzmán-Rivera, Dhruv Batra, and Pushmeet Kohli · 2012
Earlier work this paper cites.
Efficiently enforcing diversity in multi-output structured prediction
Abner Guzmán-Rivera, Pushmeet Kohli, Dhruv Batra, and Rob A. Rutenbar · 2014
Earlier work this paper cites.
Predicting multiple structured visual interpretations
Debadeepta Dey, Varun Ramakrishna, Martial Hebert, and J. Andrew Bagnell · 2015
Earlier work this paper cites.
Model predictive path integral control using covariance variable importance sampling
Grady Williams, Andrew Aldrich, and Evangelos A. Theodorou · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Earlier work this paper cites.
Variational intrinsic control
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2017
Earlier work this paper cites.
Confident multiple choice learning
Kimin Lee, Changho Hwang, KyoungSoo Park, and Jinwoo Shin · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell · 2017
Earlier work this paper cites.
Survey of model-based reinforcement learning: Applications on robotics
Athanasios S. Polydoros and Lazaros Nalpantidis · 2017
Earlier work this paper cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine · 2018
Earlier work this paper cites.
David Ha and Jürgen Schmidhuber · 2018
Earlier work this paper cites.
Soft actor-critic algorithms and applications
Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
A study on overfitting in deep reinforcement learning
Chiyuan Zhang, Oriol Vinyals, Rémi Munos, and Samy Bengio · 2018
Earlier work this paper cites.
A deep bayesian policy reuse approach against non-stationary agents
Yan Zheng, Zhaopeng Meng, Jianye Hao, Zongzhang Zhang, Tianpei Yang, and Changjie Fan · 2018
Earlier work this paper cites.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos J. Storkey, and Oleg Klimov · 2019
Earlier work this paper cites.
Diversity is all you need: Learning skills without a reward function
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine · 2019
Earlier work this paper cites.
Learning to predict without looking ahead: World models without forward prediction
C. Daniel Freeman, David Ha, and Luke Metz · 2019
Cited alongside, same era.
Learning latent dynamics for planning from pixels
Danijar Hafner, Timothy P. Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson · 2019
Cited alongside, same era.
Provably efficient maximum entropy exploration
Elad Hazan, Sham M. Kakade, Karan Singh, and Abby Van Soest · 2019
Cited alongside, same era.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Cited alongside, same era.
Plan online, learn offline: Efficient learning and exploration via model-based control
Kendall Lowrey, Aravind Rajeswaran, Sham M. Kakade, Emanuel Todorov, and Igor Mordatch · 2019
Cited alongside, same era.
Self-supervised exploration via disagreement
Deepak Pathak, Dhiraj Gandhi, and Abhinav Gupta · 2019
URLB: Unsupervised reinforcement learning benchmark
Michael Laskin, Denis Yarats, Hao Liu, Kimin Lee, Albert Zhan, Kevin Lu, Catherine Cang, Lerrel Pinto, and Pieter Abbeel · 2021
Later among the works it cites.
Model-based reinforcement learning via imagination with derived memory
Yao Mu, Yuzheng Zhuang, Bin Wang, Guangxiang Zhu, Wulong Liu, Jianyu Chen, Ping Luo, Shengbo Li, Chongjie Zhang, and Jianye Hao · 2021
Later among the works it cites.
A multi-graph attributed reinforcement learning based optimization algorithm for large-scale hybrid flow shop scheduling problem
F. Ni, J. Hao, J. Lu, X. Tong, M. Yuan, J. Duan, Y. Ma, and K. He · 2021
Later among the works it cites.
Model-based actor-critic with chance constraint for stochastic system
Baiyu Peng, Yao Mu, Yang Guan, Shengbo Eben Li, Yuming Yin, and Jianyu Chen · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Versatile multiple choice learning and its application to vision computing
Kai Tian, Yi Xu, Shuigeng Zhou, and Jihong Guan · 2019
Cited alongside, same era.
Model primitive hierarchical lifelong reinforcement learning
Bohan Wu, Jayesh K. Gupta, and Mykel J. Kochenderfer · 2019
Cited alongside, same era.
SOLAR: deep structured representations for model-based reinforcement learning
Marvin Zhang, Sharad Vikram, Laura M. Smith, Pieter Abbeel, Matthew J. Johnson, and Sergey Levine · 2019
Cited alongside, same era.
Wuji: Automatic online combat game testing using evolutionary deep reinforcement learning
Yan Zheng, Changjie Fan, Xiaofei Xie, Ting Su, Lei Ma, Jianye Hao, Zhaopeng Meng, Yang Liu, Ruimin Shen, and Yingfeng Chen · 2019
Cited alongside, same era.
Explore, discover and learn: Unsupervised discovery of state-covering skills
Victor Campos, Alexander Trott, Caiming Xiong, Richard Socher, Xavier Giró-i-Nieto, and Jordi Torres · 2020
Cited alongside, same era.
Fast task inference with variational intrinsic successor features
Steven Hansen, Will Dabney, André Barreto, David Warde-Farley, Tom Van de Wiele, and Volodymyr Mnih · 2020
Cited alongside, same era.
Aske Plaat, Walter A. Kosters, and Mike Preuss · 2021
Later among the works it cites.
Pretraining representations for data-efficient reinforcement learning
Max Schwarzer, Nitarshan Rajkumar, Michael Noukhovitch, Ankesh Anand, Laurent Charlin, R. Devon Hjelm, Philip Bachman, and Aaron C. Courville · 2021
Later among the works it cites.
State entropy maximization with random encoders for efficient exploration
Younggyo Seo, Lili Chen, Jinwoo Shin, Honglak Lee, Pieter Abbeel, and Kimin Lee · 2021
Later among the works it cites.
Learning off-policy with online planning
Harshit Sikchi, Wenxuan Zhou, and David Held · 2021
Later among the works it cites.
Unsupervised learning for reinforcement learning
Aravind Srinivas and Pieter Abbeel · 2021
Later among the works it cites.
Efficient policy detecting and reusing for non-stationarity in markov games
Yan Zheng, Jianye Hao, Zongzhang Zhang, Zhaopeng Meng, Tianpei Yang, Yanran Li, and Changjie Fan · 2021
Later among the works it cites.
Flow-based recurrent belief state learning for pomdps
Xiaoyu Chen, Yao Mark Mu, Ping Luo, Shengbo Li, and Jianyu Chen · 2022
Closest in time.
Temporal difference learning for model predictive control
Nicklas Hansen, Xiaolong Wang, and Hao Su · 2022
Closest in time.
Api: Boosting multi-agent reinforcement learning via agent-permutation-invariant networks
Xiaotian Hao, Weixun Wang, Hangyu Mao, Yaodong Yang, Dong Li, Yan Zheng, Zhen Wang, and Jianye Hao · 2022
Closest in time.
CIC: contrastive intrinsic control for unsupervised skill discovery
Michael Laskin, Hao Liu, Xue Bin Peng, Denis Yarats, Aravind Rajeswaran, and Pieter Abbeel · 2022
Closest in time.
PMIC: improving multi-agent reinforcement learning with progressive mutual information collaboration
P. Li, H. Tang, T. Yang, X. Hao, T. Sang, Y. Zheng, J. Hao, M. E. Taylor, W. Tao, and Z. Wang · 2022
Closest in time.
Good, better, best: Textual distractors generation for multiple-choice visual question answering via reinforcement learning
Jiaying Lu, Xin Ye, Yi Ren, and Yezhou Yang · 2022
Closest in time.
Curiosity-driven exploration via latent bayesian surprise
Pietro Mazzaglia, Ozan Catal, Tim Verbelen, and Bart Dhoedt · 2022
Closest in time.
Domino: Decomposed mutual information optimization for generalized context in meta-reinforcement learning
Yao Mu, Yuzheng Zhuang, Fei Ni, Bin Wang, Jianyu Chen, HAO Jianye, and Ping Luo · 2022
Closest in time.
ASE: large-scale reusable adversarial skill embeddings for physically simulated characters
Xue Bin Peng, Yunrong Guo, Lina Halper, Sergey Levine, and Sanja Fidler · 2022
Closest in time.
Unsupervised model-based pre-training for data-efficient reinforcement learning from pixels
Sai Rajeswar, Pietro Mazzaglia, Tim Verbelen, Alexandre Piché, Bart Dhoedt, Aaron Courville, and Alexandre Lacoste · 2022
Closest in time.
Reinforcement learning with action-free pre-training from videos
Younggyo Seo, Kimin Lee, Stephen L. James, and Pieter Abbeel · 2022
Closest in time.
How to leverage unlabeled data in offline reinforcement learning
Tianhe Yu, Aviral Kumar, Yevgen Chebotar, Karol Hausman, Chelsea Finn, and Sergey Levine · 2022
Closest in time.
Exploration in deep reinforcement learning: From single-agent to multiagent domain
Jianye Hao, Tianpei Yang, Hongyao Tang, Chenjia Bai, Jinyi Liu, Zhaopeng Meng, Peng Liu, and Zhen Wang · 2023
Closest in time.