Fetching the paper…
Reading the bibliography…
Learning generalist embodied agents, able to solve multitudes of tasks in different domains is a long-standing problem.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber · 2006
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Y. Bengio, N. Léonard, and A. Courville · 2013
Earlier work this paper cites.
Model-based imitation learning by probabilistic trajectory matching
P. Englert, A. Paraschos, J. Peters, and M. P. Deisenroth · 2013
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling, 2014
J. Chung, C. Gulcehre, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Concrete problems in ai safety, 2016
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané · 2016
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods, 2018
S. Fujimoto, H. van Hoof, and D. Meger · 2018
Earlier work this paper cites.
World models
D. Ha and J. Schmidhuber · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning, 2019
A. Gupta, V. Kumar, C. Lynch, S. Levine, and K. Hausman · 2019
Earlier work this paper cites.
Learning latent dynamics for planning from pixels, 2019
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson · 2019
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination, 2020
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi · 2020
Earlier work this paper cites.
Conservative q-learning for offline reinforcement learning, 2020
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems, 2020
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Earlier work this paper cites.
Planning to explore via self-supervised world models, 2020
R. Sekar, O. Rybkin, K. Daniilidis, P. Abbeel, D. Hafner, and D. Pathak · 2020
Earlier work this paper cites.
dm_control: Software and tasks for continuous control
S. Tunyasuvunakool, A. Muldal, Y. Doron, S. Liu, S. Bohez, J. Merel, T. Erez, T. Lillicrap, N. Heess, and Y. Tassa · 2020
Earlier work this paper cites.
A minimalist approach to offline reinforcement learning, 2021
S. Fujimoto and S. S. Gu · 2021
Earlier work this paper cites.
Offline reinforcement learning with implicit q-learning, 2021
I. Kostrikov, A. Nair, and S. Levine · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
Mastering visual continuous control: Improved data-augmented reinforcement learning, 2021
D. Yarats, R. Fergus, A. Lazaric, and L. Pinto · 2021
Earlier work this paper cites.
Video pretraining (vpt): Learning to act by watching unlabeled online videos, 2022
B. Baker, I. Akkaya, P. Zhokhov, J. Huizinga, J. Tang, A. Ecoffet, B. Houghton, R. Sampedro, and J. Clune · 2022
Earlier work this paper cites.
Can foundation models perform zero-shot task specification for robot manipulation?, 2022
Y. Cui, S. Niekum, A. Gupta, V. Kumar, and A. Rajeswaran · 2022
Earlier work this paper cites.
Dreamerpro: Reconstruction-free model-based reinforcement learning with prototypical representations
F. Deng, I. Jang, and S. Ahn · 2022
Earlier work this paper cites.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
L. Fan, G. Wang, Y. Jiang, A. Mandlekar, Y. Yang, H. Zhu, A. Tang, D.-A. Huang, Y. Zhu, and A. Anandkumar · 2022
Earlier work this paper cites.
Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning, 2022
W. Liang, Y. Zhang, Y. Kwon, S. Yeung, and J. Zou · 2022
Cited alongside, same era.
On the opportunities and risks of foundation models, 2022
R. Bommasani et al · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents, 2022
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Cited alongside, same era.
A generalist agent, 2022
S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Gimenez, Y. Sulsky, J. Kay, J. T. Springenberg, T. Eccles, J. Bruce, A. Razavi, A. Edwards, N. Heess, Y. Chen, R. Hadsell, O. Vinyals, M. Bordbar, and N. de Freitas · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models, 2022
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Cited alongside, same era.
Daydreamer: World models for physical robot learning, 2022
Voyager: An open-ended embodied agent with large language models, 2023
G. Wang, Y. Xie, Y. Jiang, A. Mandlekar, C. Xiao, Y. Zhu, L. Fan, and A. Anandkumar · 2023
Later among the works it cites.
Foundation models for decision making: Problems, methods, and opportunities, 2023
S. Yang, O. Nachum, Y. Du, J. Wei, P. Abbeel, and D. Schuurmans · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training, 2023
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer · 2023
Later among the works it cites.
Lumiere: A space-time diffusion model for video generation, 2024
O. Bar-Tal, H. Chefer, O. Tov, C. Herrmann, R. Paiss, S. Zada, A. Ephrat, J. Hur, G. Liu, A. Raj, Y. Li, M. Rubinstein, T. Michaeli, O. Wang, D. Sun, T. Dekel, and I. Mosseri · 2024
Closest in time.
Genie: Generative interactive environments, 2024
J. Bruce, M. Dennis, A. Edwards, J. Parker-Holder, Y. Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps, Y. Aytar, S. Bechtle, F. Behbahani, S. Chan, N. Heess, L. Gonzalez, S. Osindero, S. Ozair, S. Reed, J. Zhang, K. Zolna, J. Clune, N. de Freitas, S. Singh, and T. Rocktäschel · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Wu, A. Escontrela, D. Hafner, K. Goldberg, and P. Abbeel · 2022
Cited alongside, same era.
Don’t change the algorithm, change the data: Exploratory data for offline reinforcement learning, 2022
D. Yarats, D. Brandfonbrener, H. Liu, M. Laskin, P. Abbeel, A. Lazaric, and L. Pinto · 2022
Cited alongside, same era.
Lafite: Towards language-free training for text-to-image generation, 2022
Y. Zhou, R. Zhang, C. Chen, C. Li, C. Tensmeyer, T. Yu, J. Gu, J. Xu, and T. Sun · 2022
Cited alongside, same era.
Vision-language models as a source of rewards
K. Baumli, S. Baveja, F. Behbahani, H. Chan, G. Comanici, S. Flennerhag, M. Gazeau, K. Holsheimer, D. Horgan, M. Laskin, et al · 2023
Cited alongside, same era.
Focus: Object-centric world models for robotics manipulation, 2023
S. Ferraro, P. Mazzaglia, T. Verbelen, and B. Dhoedt · 2023
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces, 2023
A. Gu and T. Dao · 2023
Cited alongside, same era.
Mastering diverse domains through world models
D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap · 2023
Cited alongside, same era.
Simple ingredients for offline reinforcement learning, 2024
E. Cetin, A. Tirinzoni, M. Pirotta, A. Lazaric, Y. Ollivier, and A. Touati · 2024
Closest in time.
Open x-embodiment: Robotic learning datasets and rt-x models, 2024
Embodiment Collaboration et al · 2024
Closest in time.
Video prediction models as rewards for reinforcement learning
A. Escontrela, A. Adeniji, W. Yan, A. Jain, X. B. Peng, K. Goldberg, Y. Lee, D. Hafner, and P. Abbeel · 2024
Closest in time.
Gemini: A family of highly capable multimodal models, 2024
Gemini Team et al · 2024
Closest in time.
Td-mpc2: Scalable, robust world models for continuous control, 2024
N. Hansen, H. Su, and X. Wang · 2024
Closest in time.
Steve-1: A generative model for text-to-behavior in minecraft, 2024
S. Lifshitz, K. Paster, H. Chan, J. Ba, and S. McIlraith · 2024
Closest in time.
World model on million-length video and language with blockwise ringattention, 2024
H. Liu, W. Yan, M. Zaharia, and P. Abbeel · 2024
Closest in time.
Text-aware diffusion for policy learning, 2024
C. Luo, M. He, Z. Zeng, and C. Sun · 2024
Closest in time.
Eureka: Human-level reward design via coding large language models, 2024
Y. J. Ma, W. Liang, G. Wang, D.-A. Huang, O. Bastani, D. Jayaraman, Y. Zhu, L. Fan, and A. Anandkumar · 2024
Closest in time.
Gpt-4 technical report, 2024
OpenAI et al · 2024
Closest in time.
Mastering memory tasks with world models
M. R. Samsami, A. Zholus, J. Rajendran, and S. Chandar · 2024
Closest in time.
Internvideo2: Scaling video foundation models for multimodal video understanding, 2024
Y. Wang, K. Li, X. Li, J. Yu, Y. He, G. Chen, B. Pei, R. Zheng, J. Xu, Z. Wang, Y. Shi, T. Jiang, S. Li, H. Zhang, Y. Huang, Y. Qiao, Y. Wang, and L. Wang · 2024
Closest in time.
Text2reward: Reward shaping with language models for reinforcement learning, 2024
T. Xie, S. Zhao, C. H. Wu, Y. Liu, Q. Luo, V. Zhong, Y. Yang, and T. Yu · 2024
Closest in time.
Learning interactive real-world simulators, 2024
M. Yang, Y. Du, K. Ghasemipour, J. Tompson, L. Kaelbling, D. Schuurmans, and P. Abbeel · 2024
Closest in time.
Language model beats diffusion – tokenizer is key to visual generation, 2024
L. Yu, J. Lezama, N. B. Gundavarapu, L. Versari, K. Sohn, D. Minnen, Y. Cheng, V. Birodkar, A. Gupta, X. Gu, A. G. Hauptmann, B. Gong, M.-H. Yang, I. Essa, D. A. Ross, and L. Jiang · 2024
Closest in time.
Connect, collapse, corrupt: Learning cross-modal tasks with uni-modal data, 2024
Y. Zhang, E. Sui, and S. Yeung-Levy · 2024
Closest in time.
Videoprism: A foundational visual encoder for video understanding, 2024
L. Zhao, N. B. Gundavarapu, L. Yuan, H. Zhou, S. Yan, J. J. Sun, L. Friedman, R. Qian, T. Weyand, Y. Zhao, R. Hornung, F. Schroff, M.-H. Yang, D. A. Ross, H. Wang, H. Adam, M. Sirotenko, T. Liu, and B. Gong · 2024
Closest in time.