Fetching the paper…
Reading the bibliography…
Decision Transformer (DT) can learn effective policy from offline datasets by converting the offline reinforcement learning (RL) into a supervised sequence modeling task, where the trajectory elements are generated auto-regressively conditioned on the return-to-go (RTG).However, the sequence modeling learning approach tends to learn policies that converge on the sub-optimal trajectories within the dataset, for lack of bridging data to move to better trajectories, even if the condition is set to the highest RTG.To address this issue, we introduce Diffusion-Based Trajectory Branch Generation (BG), which expands the trajectories of the dataset with branches generated by a diffusion model.The trajectory branch is generated based on the segment of the trajectory within the dataset, and leads to trajectories with higher returns.We concatenate the generated branch with the trajectory segment as an expansion of the trajectory.After expanding, DT has more opportunities to learn policies to move to better trajectories, preventing it from converging to the sub-optimal trajectories.Empirically, after processing with BG, DT outperforms state-of-the-art sequence modeling methods on D4RL benchmark, demonstrating the effectiveness of adding branches to the dataset without further modifications.
Behavior regularized offline reinforcement learning
Wu, Y.; Tucker, G.; and Nachum, O. 2019 · 1911
Earlier work this paper cites.
On the estimation of production frontiers: maximum likelihood estimation of the parameters of a discontinuous density function
Aigner, D. J.; Amemiya, T.; and Poirier, D. J. 1976 · 1976
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Pomerleau, D. A. 1988 · 1988
Earlier work this paper cites.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
Siegel, N. Y.; Springenberg, J. T.; Berkenkamp, F.; Abdolmaleki, A.; Neunert, M.; Lampe, T.; Hafner, R.; Heess, N.; and Riedmiller, M. 2020 · 2002
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J.; Kumar, A.; Nachum, O.; Tucker, G.; and Levine, S. 2020 · 2004
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S.; Kumar, A.; Tucker, G.; and Fu, J. 2020 · 2005
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Song, Y.; Sohl-Dickstein, J.; Kingma, D. P.; Kumar, A.; Ermon, S.; and Poole, B. 2020 · 2011
Earlier work this paper cites.
Batch reinforcement learning
Lange, S.; Gabel, T.; and Riedmiller, M. 2012 · 2012
Earlier work this paper cites.
Geoadditive expectile regression
Sobotka, F.; and Kneib, T. 2012 · 2012
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J.; Weiss, E.; Maheswaranathan, N.; and Ganguli, S. 2015 · 2015
Earlier work this paper cites.
MIMIC-III, a freely accessible critical care database
Johnson, A. E.; Pollard, T. J.; Shen, L.; Lehman, L.-w. H.; Feng, M.; Ghassemi, M.; Moody, B.; Szolovits, P.; Anthony Celi, L.; and Mark, R. G. 2016 · 2016
Earlier work this paper cites.
1 year, 1000 km: The oxford robotcar dataset
Maddern, W.; Pascoe, G.; Linegar, C.; and Newman, P. 2017 · 2017
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T.; Zhou, A.; Abbeel, P.; and Levine, S. 2018 · 2018
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A.; Fu, J.; Soh, M.; Tucker, G.; and Levine, S. 2019 · 2019
Cited alongside, same era.
Bail: Best-action imitation learning for batch deep reinforcement learning
Chen, X.; Zhou, Z.; Wang, Z.; Wang, C.; Wu, Y.; and Ross, K. 2020 · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Kumar, A.; Zhou, A.; Tucker, G.; and Levine, S. 2020 · 2020
Cited alongside, same era.
Critic regularized regression
Wang, Z.; Novikov, A.; Zolna, K.; Merel, J. S.; Springenberg, J. T.; Reed, S. E.; Shahriari, B.; Siegel, N.; Gulcehre, C.; Heess, N.; et al. 2020 · 2020
Cited alongside, same era.
Offline rl without off-policy evaluation
Brandfonbrener, D.; Whitney, W.; Ranganath, R.; and Bruna, J. 2021 · 2021
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L.; Lu, K.; Rajeswaran, A.; Lee, K.; Grover, A.; Laskin, M.; Abbeel, P.; Srinivas, A.; and Mordatch, I. 2021 · 2021
Double check your state before trusting it: Confidence-aware bidirectional offline model-based imagination
Lyu, J.; Li, X.; and Lu, Z. 2022 · 2022
Later among the works it cites.
Online decision transformer
Zheng, Q.; Zhang, A.; and Grover, A. 2022 · 2022
Later among the works it cites.
Scalable diffusion models with transformers
Peebles, W.; and Xie, S. 2023 · 2023
Later among the works it cites.
Future-conditioned unsupervised pretraining for decision transformer
Xie, Z.; Lin, Z.; Ye, D.; Fu, Q.; Wei, Y.; and Li, S. 2023 · 2023
Later among the works it cites.
Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl
Yamagata, T.; Khalil, A.; and Santos-Rodriguez, R. 2023 · 2023
Later among the works it cites.
Uncertainty-driven trajectory truncation for data augmentation in offline reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A minimalist approach to offline reinforcement learning
Fujimoto, S.; and Gu, S. S. 2021 · 2021
Cited alongside, same era.
Offline reinforcement learning as one big sequence modeling problem
Janner, M.; Li, Q.; and Levine, S. 2021 · 2021
Cited alongside, same era.
Offline reinforcement learning with implicit q-learning
Kostrikov, I.; Nair, A.; and Levine, S. 2021 · 2021
Cited alongside, same era.
Offline reinforcement learning with reverse model-based imagination
Wang, J.; Li, W.; Jiang, H.; Zhu, G.; Li, S.; and Zhang, C. 2021 · 2021
Cited alongside, same era.
Model-based trajectory stitching for improved offline reinforcement learning
Hepburn, C. A.; and Montana, G. 2022 · 2022
Cited alongside, same era.
Elucidating the design space of diffusion-based generative models
Karras, T.; Aittala, M.; Aila, T.; and Laine, S. 2022 · 2022
Cited alongside, same era.
Zhang, J.; Lyu, J.; Ma, X.; Yan, J.; Yang, J.; Wan, L.; and Li, X. 2023 · 2023
Later among the works it cites.
Diffusion for World Modeling: Visual Details Matter in Atari
Alonso, E.; Jelley, A.; Micheli, V.; Kanervisto, A.; Storkey, A.; Pearce, T.; and Fleuret, F. 2024 · 2024
Closest in time.
Closing the Gap between TD Learning and Supervised Learning–A Generalisation Point of View
Ghugare, R.; Geist, M.; Berseth, G.; and Eysenbach, B. 2024 · 2024
Closest in time.
DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory Stitching
Li, G.; Shan, Y.; Zhu, Z.; Long, T.; and Zhang, W. 2024 · 2024
Closest in time.
Synthetic experience replay
Lu, C.; Ball, P.; Teh, Y. W.; and Parker-Holder, J. 2024 · 2024
Closest in time.
Elastic decision transformer
Wu, Y.-H.; Wang, X.; and Hamaya, M. 2024 · 2024
Closest in time.
Off-policy deep reinforcement learning without exploration
Fujimoto, S.; Meger, D.; and Precup, D. 2019 · 2062
Closest in time.