Fetching the paper…
Reading the bibliography…
Model-based reinforcement learning (MBRL) holds the promise of sample-efficient learning by utilizing a world model, which models how the environment works and typically encompasses components for two tasks: observation modeling and reward modeling.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, R. S · 1990
Earlier work this paper cites.
Multitask learning
Caruana, R · 1997
Earlier work this paper cites.
Cross-stitch networks for multi-task learning
Misra, I., Shrivastava, A., Gupta, A., and Hebert, M · 2016
Earlier work this paper cites.
Rusu, A. A., Rabinowitz, N. C., Desjardins, G., Soyer, H., Kirkpatrick, J., Kavukcuoglu, K., Pascanu, R., and Hadsell, R · 2016
Earlier work this paper cites.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K · 2017
Earlier work this paper cites.
Value prediction network
Oh, J., Singh, S., and Lee, H · 2017
Earlier work this paper cites.
An overview of multi-task learning in deep neural networks
Ruder, S · 2017
Earlier work this paper cites.
Distral: Robust multitask reinforcement learning
Teh, Y., Bapst, V., Czarnecki, W. M., Quan, J., Kirkpatrick, J., Hadsell, R., Heess, N., and Pascanu, R · 2017
Earlier work this paper cites.
Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks
Chen, Z., Badrinarayanan, V., Lee, C.-Y., and Rabinovich, A · 2018
Earlier work this paper cites.
Dynamic task prioritization for multitask learning
Guo, M., Haque, A., Huang, D.-A., Yeung, S., and Fei-Fei, L · 2018
Earlier work this paper cites.
Recurrent world models facilitate policy evolution
Ha, D. and Schmidhuber, J · 2018
Earlier work this paper cites.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
Kendall, A., Gal, Y., and Cipolla, R · 2018
Earlier work this paper cites.
Mixed precision training
Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D., Ginsburg, B., Houston, M., Kuchaiev, O., Venkatesh, G., and Wu, H · 2018
Earlier work this paper cites.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al · 2018
Earlier work this paper cites.
Natural environment benchmarks for reinforcement learning
Zhang, A., Wu, Y., and Pineau, J · 2018
Earlier work this paper cites.
Learning latent dynamics for planning from pixels
Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., and Davidson, J · 2019
Earlier work this paper cites.
End-to-end multi-task learning with attention
Liu, S., Johns, E., and Davison, A. J · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Earlier work this paper cites.
Shortcut learning in deep neural networks
Geirhos, R., Jacobsen, J.-H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., and Wichmann, F. A · 2020
Earlier work this paper cites.
Measuring visual generalization in continuous control from pixels
Grigsby, J. and Qi, Y · 2020
Cited alongside, same era.
The value equivalence principle for model-based reinforcement learning
Grimm, C., Barreto, A., Singh, S., and Silver, D · 2020
Cited alongside, same era.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M · 2020
Cited alongside, same era.
Rlbench: The robot learning benchmark & learning environment
James, S., Ma, Z., Arrojo, D. R., and Davison, A. J · 2020
Cited alongside, same era.
Model-based reinforcement learning for atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al · 2020
Cited alongside, same era.
Objective mismatch in model-based reinforcement learning
Minedojo: Building open-ended embodied agents with internet-scale knowledge
Fan, L., Wang, G., Jiang, Y., Mandlekar, A., Yang, Y., Zhu, H., Tang, A., Huang, D.-A., Zhu, Y., and Anandkumar, A · 2022
Later among the works it cites.
Benchmarking the spectrum of agent capabilities
Hafner, D · 2022
Later among the works it cites.
Temporal difference learning for model predictive control
Hansen, N., Wang, X., and Su, H · 2022
Later among the works it cites.
Q-attention: Enabling efficient learning for vision-based robotic manipulation
James, S. and Davison, A. J · 2022
Later among the works it cites.
Coarse-to-fine q-attention: Efficient learning for visual robotic manipulation via discretisation
James, S., Wada, K., Laidlow, T., and Davison, A. J · 2022
Later among the works it cites.
A path towards autonomous machine intelligence
LeCun, Y · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lambert, N., Amos, B., Yadan, O., and Calandra, R · 2020
Cited alongside, same era.
Curl: Contrastive unsupervised representations for reinforcement learning
Laskin, M., Srinivas, A., and Abbeel, P · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Cited alongside, same era.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al · 2020
Cited alongside, same era.
Unsupervised learning of visual features by contrasting cluster assignments
Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., and Joulin, A · 2021
Cited alongside, same era.
Mastering atari with discrete world models
Hafner, D., Lillicrap, T., Norouzi, M., and Ba, J · 2021
Cited alongside, same era.
Conflict-averse gradient descent for multi-task learning
Liu, B., Liu, X., Jin, X., Stone, P., and Liu, Q · 2021
Cited alongside, same era.
Later among the works it cites.
Multi-task learning as a bargaining game
Navon, A., Shamsian, A., Achituve, I., Maron, H., Kawaguchi, K., Chechik, G., and Fetaya, E · 2022
Later among the works it cites.
Denoised mdps: Learning world models better than the world itself
Wang, T., Du, S. S., Torralba, A., Isola, P., Zhang, A., and Tian, Y · 2022
Later among the works it cites.
Daydreamer: World models for physical robot learning
Wu, P., Escontrela, A., Hafner, D., Goldberg, K., and Abbeel, P · 2022
Later among the works it cites.
Mastering visual continuous control: Improved data-augmented reinforcement learning
Yarats, D., Fergus, R., Lazaric, A., and Pinto, L · 2022
Later among the works it cites.
Repo: Resilient model-based reinforcement learning by regularizing posterior predictability
Zhu, C., Simchowitz, M., Gadipudi, S., and Gupta, A · 2022
Later among the works it cites.
Mastering diverse domains through world models
Hafner, D., Pasukonis, J., Ba, J., and Lillicrap, T · 2023
Closest in time.
Forkmerge: Overcoming negative transfer in multi-task learning
Jiang, J., Chen, B., Pan, J., Wang, X., Dapeng, L., Jiang, J., and Long, M · 2023
Closest in time.
Transformers are sample efficient world models
Micheli, V., Alonso, E., and Fleuret, F · 2023
Closest in time.
Model-based reinforcement learning: A survey
Moerland, T. M., Broekens, J., Plaat, A., and Jonker, C. M · 2023
Closest in time.
Transformer-based world models are happy with 100k interactions
Robine, J., Höftmann, M., Uelwer, T., and Harmeling, S · 2023
Closest in time.
Pre-training contextualized world models with in-the-wild videos for reinforcement learning
Wu, J., Ma, H., Deng, C., and Long, M · 2023
Closest in time.
Learning from visual observation via offline pretrained state-to-go transformer
Zhou, B., Li, K., Jiang, J., and Lu, Z · 2023
Closest in time.