Fetching the paper…
Reading the bibliography…
Model-based reinforcement learning agents utilizing transformers have shown improved sample efficiency due to their ability to model extended context, resulting in more accurate world models.
Dyna, an integrated architecture for learning, planning, and reacting
Sutton, R. S · 1991
Earlier work this paper cites.
A comparison of direct and model-based reinforcement learning
Atkeson, C. G. and Santamaria, J. C · 1997
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P., Littman, M. L., and Cassandra, A. R · 1998
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Pascanu, R., Mikolov, T., and Bengio, Y · 2013
Earlier work this paper cites.
Survey of model-based reinforcement learning: Applications on robotics
Polydoros, A. S. and Nalpantidis, L · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
The dreaming variational autoencoder for reinforcement learning environments
Andersen, P.-A., Goodwin, M., and Granmo, O.-C · 2018
Earlier work this paper cites.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Buckman, J., Hafner, D., Tucker, G., Brevdo, E., and Lee, H · 2018
Earlier work this paper cites.
Efficient model-based deep reinforcement learning with variational state tabulation
Corneil, D., Gerstner, W., and Brea, J · 2018
Earlier work this paper cites.
Probabilistic recurrent state-space models
Doerr, A., Daniel, C., Schiegg, M., Duy, N.-T., Schaal, S., Toussaint, M., and Sebastian, T · 2018
Earlier work this paper cites.
Ha, D. and Schmidhuber, J · 2018
Earlier work this paper cites.
Modeling the long term future in model-based reinforcement learning
Ke, N. R., Singh, A., Touati, A., Goyal, A., Bengio, Y., Parikh, D., and Batra, D · 2018
Earlier work this paper cites.
Model-ensemble trust-region policy optimization
Kurutach, T., Clavera, I., Duan, Y., Tamar, A., and Abbeel, P · 2018
Earlier work this paper cites.
Towards sample efficient reinforcement learning
Yu, Y · 2018
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J. G., Le, Q. V., and Salakhutdinov, R · 2019
Earlier work this paper cites.
Minerl: A large-scale dataset of minecraft demonstrations
Guss, W. H., Houghton, B., Topin, N., Wang, P., Codel, C. R., Veloso, M. M., and Salakhutdinov, R · 2019
Earlier work this paper cites.
An introduction to variational autoencoders
Kingma, D. P., Welling, M., et al · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Learning to combat compounding-error in model-based reinforcement learning
Xiao, C., Wu, Y., Ma, C., Schuurmans, D., and Müller, M · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Earlier work this paper cites.
Is a good representation sufficient for sample efficient reinforcement learning?
Du, S. S., Kakade, S. M., Wang, R., and Yang, L. F · 2020
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M · 2020
Earlier work this paper cites.
On the role of planning in model-based deep reinforcement learning
Hamrick, J. B., Friesen, A. L., Behbahani, F., Guez, A., Viola, F., Witherspoon, S., Anthony, T., Buesing, L., Veličković, P., and Weber, T · 2020
Cited alongside, same era.
Rlbench: The robot learning benchmark & learning environment
James, S., Ma, Z., Arrojo, D. R., and Davison, A. J · 2020
Cited alongside, same era.
Curl: Contrastive unsupervised representations for reinforcement learning
Laskin, M., Srinivas, A., and Abbeel, P · 2020
Cited alongside, same era.
Weakly-supervised reinforcement learning for controllable behavior
Lee, L., Eysenbach, B., Salakhutdinov, R. R., Gu, S. S., and Finn, C · 2020
Cited alongside, same era.
Stabilizing transformers for reinforcement learning
Parisotto, E., Song, F., Rae, J., Pascanu, R., Gulcehre, C., Jayakumar, S., Jaderberg, M., Kaufman, R. L., Clark, A., Noury, S., et al · 2020
Cited alongside, same era.
Transdreamer: Reinforcement learning with transformer world models
Chen, C., Wu, Y.-F., Yoon, J., and Ahn, S · 2022
Later among the works it cites.
Temporal latent bottleneck: Synthesis of fast and slow processing mechanisms in sequence learning
Didolkar, A., Gupta, K., Goyal, A., Gundavarapu, N. B., Lamb, A. M., Ke, N. R., and Bengio, Y · 2022
Later among the works it cites.
Benchmarking the spectrum of agent capabilities
Hafner, D · 2022
Later among the works it cites.
Discrete compositional representations as an abstraction for goal conditioned reinforcement learning
Islam, R., Zang, H., Goyal, A., Lamb, A. M., Kawaguchi, K., Li, X., Laroche, R., Bengio, Y., and Tachet des Combes, R · 2022
Later among the works it cites.
When to update your model: Constrained model-based reinforcement learning
Ji, T., Luo, Y., Sun, F., Jing, M., He, F., and Huang, W · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al · 2020
Cited alongside, same era.
On layer normalization in the transformer architecture
Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T · 2020
Cited alongside, same era.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2020
Cited alongside, same era.
Are transformers universal approximators of sequence-to-sequence functions?
Yun, C., Bhojanapalli, S., Rawat, A. S., Reddi, S., and Kumar, S · 2020
Cited alongside, same era.
Bridging imagination and reality for model-based deep reinforcement learning
Zhu, G., Zhang, M., Lee, H., and Zhang, C · 2020
Cited alongside, same era.
Model based reinforcement learning for atari
Łukasz Kaiser, Babaeizadeh, M., Miłos, P., Osiński, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., Mohiuddin, A., Sepassi, R., Tucker, G., and Michalewski, H · 2020
Cited alongside, same era.
Deep reinforcement learning at the edge of the statistical precipice
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A. C., and Bellemare, M · 2021
Cited alongside, same era.
Transformers in vision: A survey
Khan, S., Naseer, M., Hayat, M., Zamir, S. W., Khan, F. S., and Shah, M · 2022
Later among the works it cites.
A survey of transformers
Lin, T., Wang, Y., Liu, X., and Qiu, X · 2022
Later among the works it cites.
Sample efficient deep reinforcement learning via uncertainty estimation
Mai, V., Mani, K., and Paull, L · 2022
Later among the works it cites.
Discrete representations strengthen vision transformer robustness
Mao, C., Jiang, L., Dehghani, M., Vondrick, C., Sukthankar, R., and Essa, I · 2022
Later among the works it cites.
Planning for sample efficient imitation learning
Yin, Z.-H., Ye, W., Chen, Q., and Gao, Y · 2022
Later among the works it cites.
Deep reinforcement learning with vector quantized encoding
Zhang, L., Lieffers, J., and Pyarelal, A · 2022
Later among the works it cites.
Facing off world model backbones: RNNs, transformers, and s4
Deng, F., Park, J., and Ahn, S · 2023
Later among the works it cites.
Temporal disentanglement of representations for improved generalisation in reinforcement learning
Dunion, M., McInroe, T., Luck, K. S., Hanna, J. P., and Albrecht, S. V · 2023
Later among the works it cites.
Mastering diverse domains through world models
Hafner, D., Pasukonis, J., Ba, J., and Lillicrap, T · 2023
Later among the works it cites.
Transformers are sample-efficient world models
Micheli, V., Alonso, E., and Fleuret, F · 2023
Later among the works it cites.
High-accuracy model-based reinforcement learning, a survey
Plaat, A., Kosters, W., and Preuss, M · 2023
Later among the works it cites.
Masked world models for visual control
Seo, Y., Hafner, D., Liu, H., Liu, F., James, S., Lee, K., and Abbeel, P · 2023
Later among the works it cites.
e ( 2 ) e(2) -equivariant vision transformer
Xu, R., Yang, K., Liu, K., and He, F · 2023
Later among the works it cites.
An investigation into pre-training object-centric representations for reinforcement learning
Yoon, J., Wu, Y.-F., Bae, H., and Ahn, S · 2023
Later among the works it cites.
STORM: Efficient stochastic transformer based world models for reinforcement learning
Zhang, W., Wang, G., Sun, J., Yuan, Y., and Huang, G · 2023
Later among the works it cites.
Vision transformers need registers
Darcet, T., Oquab, M., Mairal, J., and Bojanowski, P · 2024
Closest in time.
When do transformers shine in RL? decoupling memory from credit assignment
Ni, T., Ma, M., Eysenbach, B., and Bacon, P.-L · 2024
Closest in time.
Is next token prediction sufficient for gpt? exploration on code logic comprehension
Qi, M., Huang, Y., Yao, Y., Wang, M., Gu, B., and Sundaresan, N · 2024
Closest in time.