Fetching the paper…
Reading the bibliography…
Sequential modeling has demonstrated remarkable capabilities in offline reinforcement learning (RL), with Decision Transformer (DT) being one of the most notable representatives, achieving significant success.
M. Bain and C. Sammut, “A framework for behavioural cloning.” in
1995
Earlier work this paper cites.
——, “The reinforcement learning problem,”
1998
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, “Reinforcement learning: An introduction,”
1999
Earlier work this paper cites.
2013
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,”
2014
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,”
2017
Earlier work this paper cites.
Y. Ding, C. Florensa, P. Abbeel, and M. Phielipp, “Goal-conditioned imitation learning,”
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Kumar, X. B. Peng, and S. Levine, “Reward-conditioned policies,”
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
S. Levine, A. Kumar, G. Tucker, and J. Fu, “Offline reinforcement learning: Tutorial, review,”
2020
Earlier work this paper cites.
A. Kumar, A. Zhou, G. Tucker, and S. Levine, “Conservative q-learning for offline reinforcement learning,”
2020
Earlier work this paper cites.
R. Agarwal, D. Schuurmans, and M. Norouzi, “An optimistic perspective on offline reinforcement learning,” in
2020
Earlier work this paper cites.
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Y. Zou, S. Levine, C. Finn, and T. Ma, “Mopo: Model-based offline policy optimization,”
2020
Earlier work this paper cites.
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims, “Morel: Model-based offline reinforcement learning,”
2020
Earlier work this paper cites.
A. Argenson and G. Dulac-Arnold, “Model-based offline planning,”
2020
Earlier work this paper cites.
C. Lynch, M. Khansari, T. Xiao, V. Kumar, J. Tompson, S. Levine, and P. Sermanet, “Learning latent plans from play,” in
2020
Earlier work this paper cites.
A. Gu, T. Dao, S. Ermon, A. Rudra, and C. Ré, “Hippo: Recurrent memory with optimal polynomial projections,”
2020
Earlier work this paper cites.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch, “Decision transformer: Reinforcement learning via sequence modeling,”
2021
Cited alongside, same era.
A. Gu, K. Goel, and C. Re, “Efficiently modeling long sequences with structured state spaces,” in
2021
Cited alongside, same era.
M. Shridhar, L. Manuelli, and D. Fox, “Perceiver-actor: A multi-task transformer for robotic manipulation,” in
2023
Later among the works it cites.
A. Gu and T. Dao, “Mamba: Linear-time sequence modeling with selective state spaces,”
2023
Later among the works it cites.
J. Kim, S. Lee, W. Kim, and Y. Sung, “Decision convformer: Local filtering in metaformer is sufficient for decision making,” in
2023
Later among the works it cites.
T. Dao, “Flashattention-2: Faster attention with better parallelism and work partitioning,”
2023
Later among the works it cites.
B. Peng, E. Alcaide, Q. Anthony, A. Albalak, S. Arcadinho, S. Biderman, H. Cao, X. Cheng, M. Chung, L. Derczynski
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
P. Swazinna, S. Udluft, and T. Runkler, “Overcoming model bias for robust offline deep reinforcement learning,”
2021
Cited alongside, same era.
S. Emmons, B. Eysenbach, I. Kostrikov, and S. Levine, “Rvs: What is essential for offline rl via supervised learning?”
2021
Cited alongside, same era.
A. Gu, K. Goel, and C. Re, “Efficiently modeling long sequences with structured state spaces,” in
2021
Cited alongside, same era.
K. Han, A. Xiao, E. Wu, J. Guo, C. Xu, and Y. Wang, “Transformer in transformer,”
2021
Cited alongside, same era.
W. Ye, S. Liu, T. Kurutach, P. Abbeel, and Y. Gao, “Mastering atari games with limited data,”
2021
Cited alongside, same era.
S. Fujimoto and S. S. Gu, “A minimalist approach to offline reinforcement learning,”
2021
Cited alongside, same era.
I. Kostrikov, A. Nair, and S. Levine, “Offline reinforcement learning with implicit q-learning,” in
2021
Cited alongside, same era.
2023
Later among the works it cites.
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann
2023
Later among the works it cites.
L. Gao, J. Schulman, and J. Hilton, “Scaling laws for reward model overoptimization,” in
2023
Later among the works it cites.
S. Lifshitz, K. Paster, H. Chan, J. Ba, and S. McIlraith, “Steve-1: A generative model for text-to-behavior in minecraft,”
2024
Closest in time.
T. Ota, “Decision mamba: Reinforcement learning via sequence modeling with selective state spaces,”
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
M. Pióro, K. Ciebiera, K. Król, J. Ludziejewski, and S. Jaszczur, “Moe-mamba: Efficient selective state space models with mixture of experts,”
2024
Closest in time.
2024
Closest in time.
J. Liu, H. Yang, H.-Y. Zhou, Y. Xi, L. Yu, Y. Yu, Y. Liang, G. Shi, S. Zhang, H. Zheng
2024
Closest in time.
2024
Closest in time.
P. Bhargava, R. Chitnis, A. Geramifard, S. Sodhani, and A. Zhang, “When should we prefer decision transformers for offline reinforcement learning?” in
2024
Closest in time.
S. Fujimoto, D. Meger, and D. Precup, “Off-policy deep reinforcement learning without exploration,” in
2062
Closest in time.