Fetching the paper…
Reading the bibliography…
We investigate the use of transformer sequence models as dynamics models (TDMs) for control.
Problems of identification and control
K. J. Åström and B. Wittenmark · 1971
Earlier work this paper cites.
Model predictive heuristic control: Applications to industrial processes
J. Richalet, A. Rault, J. Testud, and J. Papon · 1978
Earlier work this paper cites.
Model predictive control: Theory and practice—a survey
C. E. Garcia, D. M. Prett, and M. Morari · 1989
Earlier work this paper cites.
Identification and control—closed-loop issues
P. M. Van Den Hof and R. J. Schrama · 1995
Earlier work this paper cites.
System Identification: Theory for the User
L. Ljung · 1999
Earlier work this paper cites.
The graph neural network model
F. Scarselli, M. Gori, A. C. Tsoi, M. Hagenbuchner, and G. Monfardini · 2008
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
N. Heess, G. Wayne, D. Silver, T. Lillicrap, T. Erez, and Y. Tassa · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
M. Watter, J. Springenberg, J. Boedecker, and M. Riedmiller · 2015
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Maximum a posteriori policy optimisation
A. Abdolmaleki, J. T. Springenberg, Y. Tassa, R. Munos, N. Heess, and M. Riedmiller · 2018
Earlier work this paper cites.
Relational inductive biases, deep learning, and graph networks
P. W. Battaglia, J. B. Hamrick, V. Bapst, A. Sanchez-Gonzalez, V. Zambaldi, M. Malinowski, A. Tacchetti, D. Raposo, A. Santoro, R. Faulkner, et al · 2018
Earlier work this paper cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
K. Chua, R. Calandra, R. McAllister, and S. Levine · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Graph networks as learnable physics engines for inference and control
A. Sanchez-Gonzalez, N. Heess, J. T. Springenberg, J. Merel, M. Riedmiller, R. Hadsell, and P. Battaglia · 2018
Earlier work this paper cites.
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. d. L. Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, et al · 2018
Earlier work this paper cites.
Nervenet: Learning structured policy with graph neural networks
T. Wang, R. Liao, J. Ba, and S. Fidler · 2018
Earlier work this paper cites.
Deepmdp: Learning continuous latent space models for representation learning
C. Gelada, S. Kumar, J. Buckman, O. Nachum, and M. G. Bellemare · 2019
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi · 2019
Cited alongside, same era.
Model-based reinforcement learning for atari
L. Kaiser, M. Babaeizadeh, P. Milos, B. Osinski, R. H. Campbell, K. Czechowski, D. Erhan, C. Finn, P. Kozakowski, S. Levine, et al · 2019
Cited alongside, same era.
Learning dexterous in-hand manipulation
O. M. Andrychowicz, B. Baker, M. Chociej, R. Jozefowicz, B. McGrew, J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, et al · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2022
Later among the works it cites.
Pink noise is all you need: Colored noise exploration in deep reinforcement learning
O. Eberhard, J. Hollenstein, C. Pinneri, and G. Martius · 2022
Later among the works it cites.
A system for morphology-task generalization via unified representation and behavior distillation
H. Furuta, Y. Iwasawa, Y. Matsuo, and S. S. Gu · 2022
Later among the works it cites.
Metamorph: Learning universal controllers with transformers
A. Gupta, L. Fan, S. Ganguli, and L. Fei-Fei · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Hafner, T. Lillicrap, M. Norouzi, and J. Ba · 2020
Cited alongside, same era.
One policy to control them all: Shared modular policies for agent-agnostic control
W. Huang, I. Mordatch, and D. Pathak · 2020
Cited alongside, same era.
My body is a cage: the role of morphology in graph-based incompatible control
V. Kurin, M. Igl, T. Rocktäschel, W. Boehmer, and S. Whiteson · 2020
Cited alongside, same era.
Stabilizing transformers for reinforcement learning
E. Parisotto, F. Song, J. Rae, R. Pascanu, C. Gulcehre, S. Jayakumar, M. Jaderberg, R. L. Kaufman, A. Clark, S. Noury, et al · 2020
Cited alongside, same era.
Mastering atari, go, chess and shogi by planning with a learned model
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, et al · 2020
Cited alongside, same era.
dm_control: Software and tasks for continuous control
S. Tunyasuvunakool, A. Muldal, Y. Doron, S. Liu, S. Bohez, J. Merel, T. Erez, T. Lillicrap, N. Heess, and Y. Tassa · 2020
Cited alongside, same era.
Snowflake: Scaling gnns to high-dimensional continuous control via parameter freezing
C. Blake, V. Kurin, M. Igl, and S. Whiteson · 2021
Cited alongside, same era.
Evaluating model-based planning and planner amortization for continuous control
A. Byravan, L. Hasenclever, P. Trochim, M. Mirza, A. D. Ialongo, Y. Tassa, J. T. Springenberg, A. Abdolmaleki, N. Heess, J. Merel, et al · 2021
Cited alongside, same era.
Z. Jiang, T. Zhang, M. Janner, Y. Li, T. Rocktäschel, E. Grefenstette, and Y. Tian · 2022
Later among the works it cites.
S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Gimenez, Y. Sulsky, J. Kay, J. T. Springenberg, et al · 2022
Later among the works it cites.
Planning for sample efficient imitation learning
Z.-H. Yin, W. Ye, Q. Chen, and Y. Gao · 2022
Later among the works it cites.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, W. Huang, Y. Chebotar, P. Sermanet, D. Duckworth, S. Levine, V. Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. Florence · 2023
Closest in time.
Learning agile soccer skills for a bipedal robot with deep reinforcement learning
T. Haarnoja, B. Moran, G. Lever, S. H. Huang, D. Tirumala, M. Wulfmeier, J. Humplik, S. Tunyasuvunakool, N. Y. Siegel, R. Hafner, et al · 2023
Closest in time.
Grounded decoding: Guiding text generation with grounded models for robot control
W. Huang, F. Xia, D. Shah, D. Driess, A. Zeng, Y. Lu, P. Florence, I. Mordatch, S. Levine, K. Hausman, et al · 2023
Closest in time.
Transformers are sample efficient world models
V. Micheli, E. Alonso, and F. Fleuret · 2023
Closest in time.
Model-based reinforcement learning: A survey
T. M. Moerland, J. Broekens, A. Plaat, C. M. Jonker, et al · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Predictable mdp abstraction for unsupervised model-based rl
S. Park and S. Levine · 2023
Closest in time.
Transformer-based world models are happy with 100k interactions
J. Robine, M. Höftmann, T. Uelwer, and S. Harmeling · 2023
Closest in time.
Smart: Self-supervised multi-task pretraining with control transformers
Y. Sun, S. Ma, R. Madaan, R. Bonatti, F. Huang, and A. Kapoor · 2023
Closest in time.
Foundation models for decision making: Problems, methods, and opportunities
S. Yang, O. Nachum, Y. Du, J. Wei, P. Abbeel, and D. Schuurmans · 2023
Closest in time.
Leveraging jumpy models for planning and fast learning in robotic domains
J. Zhang, J. T. Springenberg, A. Byravan, L. Hasenclever, A. Abdolmaleki, D. Rao, N. Heess, and M. Riedmiller · 2023
Closest in time.