Fetching the paper…
Reading the bibliography…
Temporal abstraction and efficient planning pose significant challenges in offline reinforcement learning, mainly when dealing with domains that involve temporally extended tasks and delayed sparse rewards.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton · 1990
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Learning options in reinforcement learning
Martin Stolle and Doina Precup · 2002
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Earlier work this paper cites.
Relative entropy policy search
Jan Peters, Katharina Mulling, and Yasemin Altun · 2010
Earlier work this paper cites.
Autonomic discovery of subgoals in hierarchical reinforcement learning
XIAO Ding, Yi-tong LI, and SHI Chuan · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Earlier work this paper cites.
Learning to see by moving
Pulkit Agrawal, Joao Carreira, and Jitendra Malik · 2015
Earlier work this paper cites.
Toward minimax off-policy value estimation
Lihong Li, Rémi Munos, and Csaba Szepesvári · 2015
Earlier work this paper cites.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
K-means clustering based reinforcement learning algorithm for automatic control in robots
Xiaobo Guo and Yan Zhai · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
A kernelized stein discrepancy for goodness-of-fit tests
Qiang Liu, Jason Lee, and Michael Jordan · 2016
Earlier work this paper cites.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba · 2017
Earlier work this paper cites.
The option-critic architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Self-consistent trajectory autoencoder: Hierarchical reinforcement learning with trajectory embeddings
John Co-Reyes, YuXuan Liu, Abhishek Gupta, Benjamin Eysenbach, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Sergey Levine · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Earlier work this paper cites.
Deep reinforcement learning and the deadly triad
Hado Van Hasselt, Yotam Doron, Florian Strub, Matteo Hessel, Nicolas Sonnerat, and Joseph Modayil · 2018
Earlier work this paper cites.
Group normalization
Yuxin Wu and Kaiming He · 2018
Earlier work this paper cites.
Search on the replay buffer: Bridging planning and reinforcement learning
Ben Eysenbach, Russ R Salakhutdinov, and Sergey Levine · 2019
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Earlier work this paper cites.
Learning latent dynamics for planning from pixels
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson · 2019
Earlier work this paper cites.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Earlier work this paper cites.
CompILE: Compositional imitation learning and execution
Thomas Kipf, Yujia Li, Hanjun Dai, Vinicius Zambaldi, Alvaro Sanchez-Gonzalez, Edward Grefenstette, Pushmeet Kohli, and Peter Battaglia · 2019
Earlier work this paper cites.
Neural probabilistic motor primitives for humanoid control
J Merel, L Hasenclever, A Galashov, A Ahuja, V Pham, G Wayne, Y Teh, and N Heess · 2019
Earlier work this paper cites.
Mish: A self regularized non-monotonic neural activation function
Diganta Misra · 2019
Earlier work this paper cites.
Learning from trajectories via subgoal discovery
Sujoy Paul, Jeroen Vanbaar, and Amit Roy-Chowdhury · 2019
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners, 2019
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Reinforcement learning upside down: Don’t predict rewards–just map them to actions
Juergen Schmidhuber · 2019
Earlier work this paper cites.
Behavior regularized offline reinforcement learning
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Earlier work this paper cites.
Hierarchical reinforcement learning for pedagogical policy induction
Guojing Zhou, Hamoon Azizsoltani, Markel Sanz Ausin, Tiffany Barnes, and Min Chi · 2019
Earlier work this paper cites.
D4RL: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Hindsight planner
Yaqing Lai, Wufan Wang, Yunjie Yang, Jihong Zhu, and Minchi Kuang · 2020
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Cited alongside, same era.
Gochat: Goal-oriented chatbots with hierarchical reinforcement learning
Jianfeng Liu, Feiyang Pan, and Ling Luo · 2020
Cited alongside, same era.
Learning latent plans from play
Corey Lynch, Mohi Khansari, Ted Xiao, Vikash Kumar, Jonathan Tompson, Sergey Levine, and Pierre Sermanet · 2020
Cited alongside, same era.
Iris: Implicit reinforcement without interaction at scale for learning control from offline robot manipulation data
Ajay Mandlekar, Fabio Ramos, Byron Boots, Silvio Savarese, Li Fei-Fei, Animesh Garg, and Dieter Fox · 2020
Cited alongside, same era.
Awac: Accelerating online reinforcement learning with offline datasets
Ashvin Nair, Abhishek Gupta, Murtaza Dalal, and Sergey Levine · 2020
Offline reinforcement learning with implicit q-learning
Ilya Kostrikov, Ashvin Nair, and Sergey Levine · 2022
Later among the works it cites.
Learning dynamic manipulation skills from haptic-play
Taeyoon Lee, Donghyun Sung, Kyoungyeon Choi, Choongin Lee, Changwoo Park, and Keunjun Choi · 2022
Later among the works it cites.
Hierarchical planning through goal-conditioned offline reinforcement learning
Jinning Li, Chen Tang, Masayoshi Tomizuka, and Wei Zhan · 2022
Later among the works it cites.
Safe offline reinforcement learning through hierarchical policies
Shaofan Liu and Shiliang Sun · 2022
Later among the works it cites.
Learning latent actions to control assistive robots
Dylan P Losey, Hong Jun Jeon, Mengxi Li, Krishnan Srinivasan, Ajay Mandlekar, Animesh Garg, Jeannette Bohg, and Dorsa Sadigh · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Hierarchical reinforcement learning with integrated discovery of salient subgoals
Shubham Pateria, Budhitama Subagdja, and Ah Hwee Tan · 2020
Cited alongside, same era.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Cited alongside, same era.
Opal: Offline primitive discovery for accelerating offline reinforcement learning
Anurag Ajay, Aviral Kumar, Pulkit Agrawal, Sergey Levine, and Ofir Nachum · 2021
Cited alongside, same era.
Laser: Learning a latent action space for efficient reinforcement learning
Arthur Allshire, Roberto Martín-Martín, Charles Lin, Shawn Manuel, Silvio Savarese, and Animesh Garg · 2021
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch · 2021
Cited alongside, same era.
A minimalist approach to offline reinforcement learning
Scott Fujimoto and Shixiang Shane Gu · 2021
Cited alongside, same era.
Mastering atari with discrete world models
Danijar Hafner, Timothy P Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2021
Cited alongside, same era.
Offline goal-conditioned reinforcement learning via f f -advantage regression
Yecheng Jason Ma, Jason Yan, Dinesh Jayaraman, and Osbert Bastani · 2022
Later among the works it cites.
Augmenting reinforcement learning with behavior primitives for diverse manipulation tasks
Soroush Nasiriany, Huihan Liu, and Yuke Zhu · 2022
Later among the works it cites.
Ase: Large-scale reusable adversarial skill embeddings for physically simulated characters
Xue Bin Peng, Yunrong Guo, Lina Halper, Sergey Levine, and Sanja Fidler · 2022
Later among the works it cites.
Learning transferable motor skills with hierarchical latent mixture policies
Dushyant Rao, Fereshteh Sadeghi, Leonard Hasenclever, Markus Wulfmeier, Martina Zambelli, Giulia Vezzani, Dhruva Tirumala, Yusuf Aytar, Josh Merel, Nicolas Heess, et al · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Later among the works it cites.
Latent plans for task-agnostic offline reinforcement learning
Erick Rosete-Beas, Oier Mees, Gabriel Kalweit, Joschka Boedecker, and Wolfram Burgard · 2022
Later among the works it cites.
Mo2: Model-based offline options
Sasha Salter, Markus Wulfmeier, Dhruva Tirumala, Nicolas Heess, Martin Riedmiller, Raia Hadsell, and Dushyant Rao · 2022
Later among the works it cites.
Bayesian nonparametrics for offline skill discovery
Valentin Villecroze, Harry Braviner, Panteha Naderian, Chris Maddison, and Gabriel Loaiza-Ganem · 2022
Later among the works it cites.
Egsde: Unpaired image-to-image translation via energy-guided stochastic differential equations
Min Zhao, Fan Bao, Chongxuan Li, and Jun Zhu · 2022
Later among the works it cites.
Online decision transformer
Qinqing Zheng, Amy Zhang, and Aditya Grover · 2022
Later among the works it cites.
Diffusion policies for out-of-distribution generalization in offline reinforcement learning
Suzan Ece Ada, Erhan Oztop, and Emre Ugur · 2023
Closest in time.
Is conditional generative modeling all you need for decision-making?
Anurag Ajay, Yilun Du, Abhi Gupta, Joshua B Tenenbaum, Tommi S Jaakkola, and Pulkit Agrawal · 2023
Closest in time.
Equivariant energy-guided SDE for inverse molecular design
Fan Bao, Min Zhao, Zhongkai Hao, Peiyao Li, Chongxuan Li, and Jun Zhu · 2023
Closest in time.
Sequence modeling is a robust contender for offline reinforcement learning
Prajjwal Bhargava, Rohan Chitnis, Alborz Geramifard, Shagun Sodhani, and Amy Zhang · 2023
Closest in time.
Imitating complex trajectories: Bridging low-level stability and high-level behavior
Adam Block, Daniel Pfrommer, and Max Simchowitz · 2023
Closest in time.
Iql-td-mpc: Implicit q-learning for hierarchical model predictive control
Rohan Chitnis, Yingchen Xu, Bobak Hashemi, Lucas Lehnert, Urun Dogan, Zheqing Zhu, and Olivier Delalleau · 2023
Closest in time.
Diffusion posterior sampling for general noisy inverse problems
Hyungjin Chung, Jeongsol Kim, Michael Thompson Mccann, Marc Louis Klasky, and Jong Chul Ye · 2023
Closest in time.
Mastering diverse domains through world models
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap · 2023
Closest in time.
Diffusion model is an effective planner and data synthesizer for multi-task reinforcement learning
Haoran He, Chenjia Bai, Kang Xu, Zhuoran Yang, Weinan Zhang, Dong Wang, Bin Zhao, and Xuelong Li · 2023
Closest in time.
Instructed diffuser with temporal condition guidance for offline reinforcement learning
Jifeng Hu, Yanchao Sun, Sili Huang, SiYuan Guo, Hechang Chen, Li Shen, Lichao Sun, Yi Chang, and Dacheng Tao · 2023
Closest in time.
Chain-of-thought predictive control
Zhiwei Jia, Fangchen Liu, Vineet Thumuluri, Linghao Chen, Zhiao Huang, and Hao Su · 2023
Closest in time.
Efficient planning in a compact latent action space
Zhengyao Jiang, Tianjun Zhang, Michael Janner, Yueying Li, Tim Rocktäschel, Edward Grefenstette, and Yuandong Tian · 2023
Closest in time.
Hierarchical imitation learning with vector quantized models
Kalle Kujanpää, Joni Pajarinen, and Alexander Ilin · 2023
Closest in time.
Hierarchical diffusion for offline decision making
Wenhao Li, Xiangfeng Wang, Bo Jin, and Hongyuan Zha · 2023
Closest in time.
Adaptdiffuser: Diffusion models as adaptive self-evolving planners
Zhixuan Liang, Yao Mu, Mingyu Ding, Fei Ni, Masayoshi Tomizuka, and Ping Luo · 2023
Closest in time.
Metadiffuser: Diffusion model as conditional planner for offline meta-rl
Fei Ni, Jianye Hao, Yao Mu, Yifu Yuan, Yan Zheng, Bin Wang, and Zhixuan Liang · 2023
Closest in time.
Imitating human behaviour with diffusion models
Tim Pearce, Tabish Rashid, Anssi Kanervisto, Dave Bignell, Mingfei Sun, Raluca Georgescu, Sergio Valcarcel Macua, Shan Zheng Tan, Ida Momennejad, Katja Hofmann, and Sam Devlin · 2023
Closest in time.
Dreamfusion: Text-to-3d using 2d diffusion
Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Mildenhall · 2023
Closest in time.
Consistency models
Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever · 2023
Closest in time.
Fighting uncertainty with gradients: Offline reinforcement learning via diffusion score matching
HJ Suh, Glen Chou, Hongkai Dai, Lujie Yang, Abhishek Gupta, and Russ Tedrake · 2023
Closest in time.
Latent skill models for offline reinforcement learning
Siddarth Venkatraman · 2023
Closest in time.
Diffusion policies as an expressive policy class for offline reinforcement learning
Zhendong Wang, Jonathan J Hunt, and Mingyuan Zhou · 2023
Closest in time.
Yueh-Hua Wu, Xiaolong Wang, and Masashi Hamaya · 2023
Closest in time.
Flow to control: Offline reinforcement learning with lossless primitive discovery
Yiqin Yang, Hao Hu, Wenzhe Li, Siyuan Li, Jun Yang, Qianchuan Zhao, and Chongjie Zhang · 2023
Closest in time.