Fetching the paper…
Reading the bibliography…
Offline Goal-Conditioned RL (GCRL) offers a feasible paradigm for learning general-purpose policies from diverse and multi-task offline datasets.
Efficient selectivity and backup operators in monte-carlo tree search
Rémi Coulom · 2007
Earlier work this paper cites.
PILCO: A model-based and data-efficient approach to policy search
Marc Deisenroth and Carl E Rasmussen · 2011
Earlier work this paper cites.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel, and Wojciech Zaremba · 2017
Earlier work this paper cites.
Multi-goal reinforcement learning: Challenging robotics environments and request for research, 2018
Matthias Plappert, Marcin Andrychowicz, Alex Ray, Bob McGrew, Bowen Baker, Glenn Powell, Jonas Schneider, Josh Tobin, Maciek Chociej, Peter Welinder, Vikash Kumar, and Wojciech Zaremba · 2018
Earlier work this paper cites.
Exponentially weighted imitation learning for batched historical data
Qing Wang, Jiechao Xiong, Lei Han, peng sun, Han Liu, and Tong Zhang · 2018
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Earlier work this paper cites.
Learning latent dynamics for planning from pixels
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson · 2019
Earlier work this paper cites.
Learning latent plans from play
Corey Lynch, Mohi Khansari, Ted Xiao, Vikash Kumar, Jonathan Tompson, Sergey Levine, and Pierre Sermanet · 2019
Earlier work this paper cites.
Self-supervised exploration via disagreement
Deepak Pathak, Dhiraj Gandhi, and Abhinav Gupta · 2019
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning, 2019
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2019
Earlier work this paper cites.
A perspective on objects and systematic generalization in model-based RL, 2019
Sjoerd van Steenkiste, Klaus Greff, and Jürgen Schmidhuber · 2019
Earlier work this paper cites.
Behavior regularized offline reinforcement learning, 2019
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Earlier work this paper cites.
PlanGAN: Model-based planning with sparse rewards and multiple goals
Henry Charlesworth and Giovanni Montana · 2020
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi · 2020
Earlier work this paper cites.
MOReL: Model-based offline reinforcement learning
Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims · 2020
Earlier work this paper cites.
Conservative Q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Earlier work this paper cites.
Context-aware dynamics model for generalization in model-based reinforcement learning
Kimin Lee, Younggyo Seo, Seunghyun Lee, Honglak Lee, and Jinwoo Shin · 2020
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems, 2020
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Earlier work this paper cites.
Learning latent plans from play
Corey Lynch, Mohi Khansari, Ted Xiao, Vikash Kumar, Jonathan Tompson, Sergey Levine, and Pierre Sermanet · 2020
Cited alongside, same era.
Deep dynamics models for learning dexterous manipulation
Anusha Nagabandi, Kurt Konolige, Sergey Levine, and Vikash Kumar · 2020
Cited alongside, same era.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, Timothy Lillicrap, and David Silver · 2020
Cited alongside, same era.
MOPO: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Y Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Cited alongside, same era.
PLAS: Latent action space for offline reinforcement learning, 2020
Wenxuan Zhou, Sujay Bajracharya, and David Held · 2020
Cited alongside, same era.
Uncertainty-based offline reinforcement learning with diversified q-ensemble
LAPO: Latent-variable advantage-weighted policy optimization for offline reinforcement learning
Xi Chen, Ali Ghadirzadeh, Tianhe Yu, Jianhao Wang, Alex Yuan Gao, Wenzhe Li, Liang Bin, Chelsea Finn, and Chongjie Zhang · 2022
Later among the works it cites.
RvS: What is essential for offline RL via supervised learning?
Scott Emmons, Benjamin Eysenbach, Ilya Kostrikov, and Sergey Levine · 2022
Later among the works it cites.
Contrastive learning as goal-conditioned reinforcement learning
Benjamin Eysenbach, Tianjun Zhang, Sergey Levine, and Ruslan Salakhutdinov · 2022
Later among the works it cites.
Temporal difference learning for model predictive control
Nicklas A. Hansen, Hao Su, and Xiaolong Wang · 2022
Later among the works it cites.
Planning with diffusion for flexible behavior synthesis, 2022
Michael Janner, Yilun Du, Joshua B. Tenenbaum, and Sergey Levine · 2022
Later among the works it cites.
Offline goal-conditioned reinforcement learning via f-advantage regression
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gaon An, Seungyong Moon, Jang-Hyun Kim, and Hyun Oh Song · 2021
Cited alongside, same era.
Model-based offline planning
Arthur Argenson and Gabriel Dulac-Arnold · 2021
Cited alongside, same era.
Actionable models: Unsupervised offline reinforcement learning of robotic skills
Yevgen Chebotar, Karol Hausman, Yao Lu, Ted Xiao, Dmitry Kalashnikov, Jacob Varley, Alex Irpan, Benjamin Eysenbach, Ryan C Julian, Chelsea Finn, and Sergey Levine · 2021
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch · 2021
Cited alongside, same era.
A minimalist approach to offline reinforcement learning
Scott Fujimoto and Shixiang Gu · 2021
Cited alongside, same era.
Learning to reach goals via iterated supervised learning
Dibya Ghosh, Abhishek Gupta, Ashwin Reddy, Justin Fu, Coline Manon Devin, Benjamin Eysenbach, and Sergey Levine · 2021
Cited alongside, same era.
Offline reinforcement learning as one big sequence modeling problem
Michael Janner, Qiyang Li, and Sergey Levine · 2021
Cited alongside, same era.
Yecheng Jason Ma, Jason Yan, Dinesh Jayaraman, and Osbert Bastani · 2022
Later among the works it cites.
Learning goal-conditioned policies offline with self-supervised reward shaping
Lina Mezghani, Sainbayar Sukhbaatar, Piotr Bojanowski, Alessandro Lazaric, and Karteek Alahari · 2022
Later among the works it cites.
RAMBO-RL: Robust adversarial model-based offline reinforcement learning
Marc Rigter, Bruno Lacerda, and Nick Hawes · 2022
Later among the works it cites.
Can push-forward generative models fit multimodal distributions?, 2022
Antoine Salmona, Valentin de Bortoli, Julie Delon, and Agnès Desolneux · 2022
Later among the works it cites.
Behavior transformers: Cloning k modes with one stone
Nur Muhammad Mahi Shafiullah, Zichen Jeff Cui, Ariuntuya Altanzaya, and Lerrel Pinto · 2022
Later among the works it cites.
Exploit reward shifting in value-based deep-rl: Optimistic curiosity-based exploration and conservative exploitation via linear reward shaping
Hao Sun, Lei Han, Rui Yang, Xiaoteng Ma, Jian Guo, and Bolei Zhou · 2022
Later among the works it cites.
Dual generator offline reinforcement learning, 2022
Quan Vuong, Aviral Kumar, Sergey Levine, and Yevgen Chebotar · 2022
Later among the works it cites.
Model-based offline planning with trajectory pruning
Xianyuan Zhan, Xiangyu Zhu, and Haoran Xu · 2022
Later among the works it cites.
Model-based reinforcement learning: A survey
Thomas M Moerland, Joost Broekens, Aske Plaat, Catholijn M Jonker, et al · 2023
Closest in time.
Entropy: Environment transformer and offline policy optimization, 2023
Pengqin Wang, Meixin Zhu, and Shaojie Shen · 2023
Closest in time.
What is essential for unseen goal generalization of offline goal-conditioned RL?
Rui Yang, Yong Lin, Xiaoteng Ma, Hao Hu, Chongjie Zhang, and Tong Zhang · 2023
Closest in time.
Goal-conditioned offline reinforcement learning through state space partitioning
Mianchu Wang, Yue Jin, and Giovanni Montana · 2024
Closest in time.