Fetching the paper…
Reading the bibliography…
Both text and video data are abundant on the internet and support large-scale self-supervised learning through next token or frame prediction.
Minds, brains, and programs
Searle, J. R · 1980
Earlier work this paper cites.
Society of mind
Minsky, M · 1988
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Sutton, R. S · 1991
Earlier work this paper cites.
Consciousness explained
Dennett, D. C · 1993
Earlier work this paper cites.
Between MDPs and semi-MDPs: a framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Image-based computational fluid dynamics modeling in realistic arterial geometries
Steinman, D. A · 2002
Earlier work this paper cites.
If a picture is worth a thousand words is video worth a million? differences in affective and cognitive processing of video and text cases
Yadav, A., Phillips, M. M., Lundeberg, M. A., Koehler, M. J., Hilden, K., and Dirkin, K. H · 2011
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Sim-to-real robot learning from pixels with progressive nets
Rusu, A. A., Večerík, M., Rothörl, T., Heess, N., Pascanu, R., and Hadsell, R · 2017
Earlier work this paper cites.
The predictron: End-to-end learning and planning
Silver, D., van Hasselt, H., Hessel, M., Schaul, T., Guez, A., Harley, T., Dulac-Arnold, G., Reichert, D. P., Rabinowitz, N. C., Barreto, A., and Degris, T · 2017
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., and Abbeel, P · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al · 2017
Earlier work this paper cites.
A trillion frames per second: the techniques and applications of light-in-flight photography
Faccio, D. and Velten, A · 2018
Earlier work this paper cites.
Illuminating generalization in deep reinforcement learning through procedural level generation
Justesen, N., Torrado, R. R., Bontrager, P., Khalifa, A., Togelius, J., and Risi, S · 2018
Earlier work this paper cites.
Learning what you can do before doing anything
Rybkin, O., Pertsch, K., Derpanis, K. G., Daniilidis, K., and Jaegle, A · 2018
Earlier work this paper cites.
Procedural content generation via machine learning (PCGML)
Summerville, A., Snodgrass, S., Guzdial, M., Holmgård, C., Hoover, A. K., Isaksen, A., Nealen, A., and Togelius, J · 2018
Earlier work this paper cites.
Artificial Intelligence and Games
Yannakakis, G. N. and Togelius, J · 2018
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2019
Earlier work this paper cites.
Imitating latent policies from observation
Edwards, A., Sahni, H., Schroecker, Y., and Isbell, C · 2019
Earlier work this paper cites.
Robot motion planning in learned latent spaces
Ichter, B. and Pavone, M · 2019
Earlier work this paper cites.
Generating diverse high-fidelity images with vq-vae-2
Razavi, A., Van den Oord, A., and Vinyals, O · 2019
Earlier work this paper cites.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., Lillicrap, T. P., and Silver, D · 2019
Earlier work this paper cites.
Neural game engine: Accurate learning of generalizable forward models from pixels
Bamford, C. and Lucas, S. M · 2020
Earlier work this paper cites.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J · 2020
Earlier work this paper cites.
Mastering atari with discrete world models
Hafner, D., Lillicrap, T., Norouzi, M., and Ba, J · 2020
Earlier work this paper cites.
Learning to Simulate Dynamic Environments with GameGAN
Kim, S. W., Zhou, Y., Philion, J., Torralba, A., and Fidler, S · 2020
Cited alongside, same era.
Increasing generality in machine learning through procedural content generation
Risi, S. and Togelius, J · 2020
Cited alongside, same era.
A tutorial on se(3) transformation parameterizations and on-manifold optimization
Blanco-Claraco, J. L · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Cited alongside, same era.
Neural network analysis of electron microscopy video data reveals the temperature-driven microphase dynamics in the ions/water system
Kashin, A. S., Boiko, D. A., and Ananikov, V. P · 2021
Cited alongside, same era.
Video prediction models as rewards for reinforcement learning
Escontrela, A., Adeniji, A., Yan, W., Jain, A., Peng, X. B., Goldberg, K., Lee, Y., Hafner, D., and Abbeel, P · 2023
Later among the works it cites.
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Guo, Y., Yang, C., Rao, A., Wang, Y., Qiao, Y., Lin, D., and Dai, B · 2023
Later among the works it cites.
Animate-a-story: Storytelling with retrieval-augmented video generation
He, Y., Xia, M., Chen, H., Cun, X., Gong, Y., Xing, J., Zhang, Y., Wang, X., Weng, C., Shan, Y., et al · 2023
Later among the works it cites.
Let’s think frame by frame with vip: A video infilling and prediction dataset for evaluating video chain-of-thought
Himakunthala, V., Ouyang, A., Rose, D., He, R., Mei, A., Lu, Y., Sonar, C., Saxon, M., and Wang, W · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Planning in stochastic environments with a learned model
Antonoglou, I., Schrittwieser, J., Ozair, S., Hubert, T. K., and Silver, D · 2022
Cited alongside, same era.
Video pretraining (vpt): Learning to act by watching unlabeled online videos
Baker, B., Akkaya, I., Zhokov, P., Huizinga, J., Tang, J., Ecoffet, A., Houghton, B., Sampedro, R., and Clune, J · 2022
Cited alongside, same era.
Visual prompting via image inpainting
Bar, A., Gandelsman, Y., Darrell, T., Globerson, A., and Efros, A · 2022
Cited alongside, same era.
Maskgit: Masked generative image transformer
Chang, H., Zhang, H., Jiang, L., Liu, C., and Freeman, W. T · 2022
Cited alongside, same era.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
Fan, L., Wang, G., Jiang, Y., Mandlekar, A., Yang, Y., Zhu, H., Tang, A., Huang, D.-A., Zhu, Y., and Anandkumar, A · 2022
Cited alongside, same era.
Maskvit: Masked visual pre-training for video prediction
Gupta, A., Tian, S., Zhang, Y., Wu, J., Martín-Martín, R., and Fei-Fei, L · 2022
Cited alongside, same era.
Classifier-free diffusion guidance
Ho, J. and Salimans, T · 2022
Cited alongside, same era.
Hu, A., Russell, L., Yeo, H., Murez, Z., Fedoseev, G., Kendall, A., Shotton, J., and Corrado, G · 2023
Later among the works it cites.
Imagined subgoals for hierarchical goal-conditioned policies
Kang, X., Ye, W., and Kuo, Y.-L · 2023
Later among the works it cites.
Learning to act from actionless videos through dense correspondences, 2023
Ko, P.-C., Mao, J., Du, Y., Sun, S.-H., and Tenenbaum, J. B · 2023
Later among the works it cites.
Videopoet: A large language model for zero-shot video generation
Kondratyuk, D., Yu, L., Gu, X., Lezama, J., Huang, J., Hornung, R., Adam, H., Akbari, H., Alon, Y., Birodkar, V., et al · 2023
Later among the works it cites.
Li, Z., Tucker, R., Snavely, N., and Holynski, A · 2023
Later among the works it cites.
Interactive language: Talking to robots in real time
Lynch, C., Wahid, A., Tompson, J., Ding, T., Betker, J., Baruch, R., Armstrong, T., and Florence, P · 2023
Later among the works it cites.
Augmented language models: a survey
Mialon, G., Dessì, R., Lomeli, M., Nalmpantis, C., Pasunuru, R., Raileanu, R., Rozière, B., Schick, T., Dwivedi-Yu, J., Celikyilmaz, A., et al · 2023
Later among the works it cites.
Open x-embodiment: Robotic learning datasets and rt-x models
Padalkar, A., Pooley, A., Jain, A., Bewley, A., Herzog, A., Irpan, A., Khazatsky, A., Rai, A., Singh, A., Brohan, A., et al · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Later among the works it cites.
Learning silicon dopant transitions in graphene using scanning transmission electron microscopy
Schwarzer, M., Farebrother, J., Greaves, J., Roccapriore, K., Cubuk, E., Agarwal, R., Courville, A., Bellemare, M., Kalinin, S., Mordatch, I., et al · 2023
Later among the works it cites.
Genhowto: Learning to generate actions and state transformations from instructional videos
Souček, T., Damen, D., Wray, M., Laptev, I., and Sivic, J · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Later among the works it cites.
Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change
Valmeekam, K., Marquez, M., Olmo, A., Sreedharan, S., and Kambhampati, S · 2023
Later among the works it cites.
Any-point trajectory modeling for policy learning
Wen, C., Lin, X., So, J., Chen, K., Dou, Q., Gao, Y., and Abbeel, P · 2023
Later among the works it cites.
Art ⋅ \boldsymbol{\cdot} v: Auto-regressive text-to-video generation with diffusion models, 2023
Weng, W., Feng, R., Wang, Y., Dai, Q., Wang, C., Yin, D., Zhao, Z., Qiu, K., Bao, J., Yuan, Y., Luo, C., Zhang, Y., and Xiong, Z · 2023
Later among the works it cites.
A survey on video diffusion models
Xing, Z., Feng, Q., Chen, H., Dai, Q., Hu, H., Xu, H., Wu, Z., and Jiang, Y.-G · 2023
Later among the works it cites.
Temporally consistent transformers for video generation
Yan, W., Hafner, D., James, S., and Abbeel, P · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K · 2023
Later among the works it cites.
Scaling robot learning with semantically imagined experience
Yu, T., Xiao, T., Stone, A., Tompson, J., Brohan, A., Wang, S., Singh, J., Tan, C., Peralta, J., Ichter, B., et al · 2023
Later among the works it cites.
Zhu, J., Yang, H., He, H., Wang, W., Tuo, Z., Cheng, W.-H., Gao, L., Song, J., and Fu, J · 2023
Later among the works it cites.
Lumiere: A space-time diffusion model for video generation
Bar-Tal, O., Chefer, H., Tov, O., Herrmann, C., Paiss, R., Zada, S., Ephrat, A., Hur, J., Li, Y., Michaeli, T., et al · 2024
Closest in time.
Genie: Generative interactive environments, 2024
Bruce, J., Dennis, M., Edwards, A., Parker-Holder, J., Shi, Y., Hughes, E., Lai, M., Mavalankar, A., Steigerwald, R., Apps, C., Aytar, Y., Bechtle, S., Behbahani, F., Chan, S., Heess, N., Gonzalez, L., Osindero, S., Ozair, S., Reed, S., Zhang, J., Zolna, K., Clune, J., de Freitas, N., Singh, S., and Rocktäschel, T · 2024
Closest in time.
Solving olympiad geometry without human demonstrations
Trinh, T. H., Wu, Y., Le, Q. V., He, H., and Luong, T · 2024
Closest in time.