Fetching the paper…
Reading the bibliography…
Decision Transformer (DT) is a recently proposed architecture for Reinforcement Learning that frames the decision-making process as an auto-regressive sequence modeling problem and uses a Transformer model to predict the next action in a sequence of states, actions, and rewards.
Learning internal representations by error propagation
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Serial order: A parallel distributed processing approach
M. I. Jordan · 1997
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
An empirical evaluation of generic convolutional and recurrent networks for sequence modeling
S. Bai, J. Z. Kolter, and V. Koltun · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al · 2018
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Cited alongside, same era.
Reinforcement learning upside down: Don’t predict rewards - just map them to actions
J. Schmidhuber · 2019
Cited alongside, same era.
Understanding lstm – a tutorial into long short-term memory recurrent neural networks, 2019
R. C. Staudemeyer and E. R. Morris · 2019
Cited alongside, same era.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, et al · 2020
Cited alongside, same era.
D4RL: datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Cited alongside, same era.
Offline pre-trained multi-agent decision transformer: One big sequence model tackles all SMAC tasks
L. Meng, M. Wen, Y. Yang, C. Le, et al · 2021
Later among the works it cites.
Understanding the effects of dataset characteristics on offline reinforcement learning
K. Schweighofer, M. Hofmarcher, M.-C. Dinu, P. Renz, A. Bitto-Nemling, V. P. Patil, and S. Hochreiter · 2021
Later among the works it cites.
Generalized decision transformer for offline hindsight information matching
H. Furuta, Y. Matsuo, and S. S. Gu · 2022
Closest in time.
Transformers in vision: A survey
S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah · 2022
Closest in time.
Goal-conditioned reinforcement learning: Problems and solutions
M. Liu, M. Zhu, and W. Zhang · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Memory based trajectory-conditioned policies for learning from sparse rewards
Y. Guo, J. Choi, M. Moczulski, S. Feng, S. Bengio, M. Norouzi, and H. Lee · 2020
Cited alongside, same era.
Sim-to-real transfer in deep reinforcement learning for robotics: a survey
W. Zhao, J. P. Queralta, and T. Westerlund · 2020
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, et al · 2021
Cited alongside, same era.
Offline reinforcement learning as one big sequence modeling problem
M. Janner, Q. Li, and S. Levine · 2021
Cited alongside, same era.
Robot learning from randomized simulations: A review
F. Muratore, F. Ramos, G. Turk, W. Yu, M. Gienger, and J. Peters · 2022
Closest in time.
S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, et al · 2022
Closest in time.
Can wikipedia help offline reinforcement learning?
M. Reid, Y. Yamada, and S. S. Gu · 2022
Closest in time.
Online decision transformer
Q. Zheng, A. Zhang, and A. Grover · 2022
Closest in time.