Fetching the paper…
Reading the bibliography…
We introduce Masked Trajectory Models (MTM) as a generic abstraction for sequential decision making.
Dynamic programming and optimal control
Bertsekas, D. P · 1995
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Approximate dynamic programming - solving the curses of dimensionality
Powell, W. B · 2007
Earlier work this paper cites.
Feedback systems: An introduction for scientists and engineers
Åström, K. J. and Murray, R. M · 2008
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Vincent, P., Larochelle, H., Bengio, Y., and Manzagol, P.-A · 2008
Earlier work this paper cites.
Heteromodal cortex
Donnelly, K · 2011
Earlier work this paper cites.
Batch reinforcement learning
Lange, S., Gabel, T., and Riedmiller, M · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Layer normalization, 2016
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Openai gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Principal component analysis: a review and recent developments
Jolliffe, I. T. and Cadima, J · 2016
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
Kingma, D. P. and Ba, J · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations
Rajeswaran, A., Kumar, V., Gupta, A., Vezzani, G., Schulman, J., Todorov, E., and Levine, S · 2018
Earlier work this paper cites.
When to trust your model: Model-based policy optimization
Janner, M., Fu, J., Zhang, M., and Levine, S · 2019
Earlier work this paper cites.
Self-supervised visual feature learning with deep neural networks: A survey
Jing, L. and Tian, Y · 2019
Cited alongside, same era.
Roberta: A robustly optimized BERT pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Cited alongside, same era.
Decoupled weight decay regularization, 2019
Loshchilov, I. and Hutter, F · 2019
Cited alongside, same era.
Reinforcement learning upside down: Don’t predict rewards – just map them to actions, 2019
Schmidhuber, J · 2019
Cited alongside, same era.
Training agents using upside-down reinforcement learning
Srivastava, R. K., Shyam, P., Mutz, F., Jaśkowski, W., and Schmidhuber, J · 2019
Cited alongside, same era.
Urlb: Unsupervised reinforcement learning benchmark
Laskin, M., Yarats, D., Liu, H., Lee, K., Zhan, A., Lu, K., Cang, C., Pinto, L., and Abbeel, P · 2021
Later among the works it cites.
State-only imitation learning for dexterous manipulation
Radosavovic, I., Wang, X., Pinto, L., and Malik, J · 2021
Later among the works it cites.
Representation matters: Offline pretraining for sequential decision making, 2021
Yang, M. and Nachum, O · 2021
Later among the works it cites.
Reinforcement learning with prototypical representations
Yarats, D., Fergus, R., Lazaric, A., and Pinto, L · 2021
Later among the works it cites.
Combo: Conservative offline model-based policy optimization
Yu, T., Kumar, A., Rafailov, R., Rajeswaran, A., Levine, S., and Finn, C · 2021
Later among the works it cites.
Video pretraining (vpt): Learning to act by watching unlabeled online videos
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning, 2020
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Mastering atari with discrete world models
Hafner, D., Lillicrap, T., Norouzi, M., and Ba, J · 2020
Cited alongside, same era.
MOReL : Model-Based Offline Reinforcement Learning
Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T · 2020
Cited alongside, same era.
Conservative Q-Learning for Offline Reinforcement Learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Provably good batch off-policy reinforcement learning without great exploration
Liu, Y., Swaminathan, A., Agarwal, A., and Brunskill, E · 2020
Cited alongside, same era.
A Game Theoretic Framework for Model-Based Reinforcement Learning
Rajeswaran, A., Mordatch, I., and Kumar, V · 2020
Cited alongside, same era.
Baker, B., Akkaya, I., Zhokhov, P., Huizinga, J., Tang, J., Ecoffet, A., Houghton, B., Sampedro, R., and Clune, J · 2022
Later among the works it cites.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al · 2022
Later among the works it cites.
Unimask: Unified inference in sequential decision problems
Carroll, M., Paradise, O., Lin, J., Georgescu, R., Sun, M., Bignell, D., Milani, S., Hofmann, K., Hausknecht, M. J., Dragan, A. D., and Devlin, S · 2022
Later among the works it cites.
Masked autoencoding for scalable and generalizable decision making
Liu, F., Liu, H., Grover, A., and Abbeel, P · 2022
Later among the works it cites.
R3m: A universal visual representation for robot manipulation
Nair, S., Rajeswaran, A., Kumar, V., Finn, C., and Gupta, A · 2022
Later among the works it cites.
The unsurprising effectiveness of pre-trained vision models for control
Parisi, S., Rajeswaran, A., Purushwalkam, S., and Gupta, A. K · 2022
Later among the works it cites.
Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., Eccles, T., Bruce, J., Razavi, A., Edwards, A. D., Heess, N. M. O., Chen, Y., Hadsell, R., Vinyals, O., Bordbar, M., and de Freitas, N · 2022
Later among the works it cites.
Masked world models for visual control
Seo, Y., Hafner, D., Liu, H., Liu, F., James, S., Lee, K., and Abbeel, P · 2022
Later among the works it cites.
Behavior transformers: Cloning k modes with one stone
Shafiullah, N. M. M., Cui, Z. J., Altanzaya, A., and Pinto, L · 2022
Later among the works it cites.
Masked visual pre-training for motor control
Xiao, T., Radosavovic, I., Darrell, T., and Malik, J · 2022
Later among the works it cites.
Don’t change the algorithm, change the data: Exploratory data for offline reinforcement learning
Yarats, D., Brandfonbrener, D., Liu, H., Laskin, M., Abbeel, P., Lazaric, A., and Pinto, L · 2022
Later among the works it cites.
Semi-supervised offline reinforcement learning with action-free trajectories, 10 2022
Zheng, Q., Henaff, M., Amos, B., and Grover, A · 2022
Later among the works it cites.
Policy architectures for compositional generalization in control
Zhou, A., Kumar, V., Finn, C., and Rajeswaran, A · 2022
Later among the works it cites.