Fetching the paper…
Reading the bibliography…
Video generation models have made significant progress in simulating future states, showcasing their potential as world simulators in embodied scenarios.
Adaptive mixtures of local experts
Jacobs, R. A., Jordan, M. I., Nowlan, S. J., and Hinton, G. E · 1991
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Banerjee, S. and Lavie, A · 2005
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Chen, X., Fang, H., Lin, T.-Y., Vedantam, R., Gupta, S., Dollár, P., and Zitnick, C. L · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Vedantam, R., Lawrence Zitnick, C., and Parikh, D · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
Anderson, P., Fernando, B., Johnson, M., and Gould, S · 2016
Earlier work this paper cites.
Ha, D. and Schmidhuber, J · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Unterthiner, T., Van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S · 2018
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A · 2021
Earlier work this paper cites.
Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers
Chefer, H., Gur, S., and Wolf, L · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al · 2022
Earlier work this paper cites.
Ego4d: Around the world in 3,000 hours of egocentric video
Grauman, K., Westbury, A., Byrne, E., Chavis, Z., Furnari, A., Girdhar, R., Hamburger, J., Jiang, H., Liu, M., Liu, X., et al · 2022
Earlier work this paper cites.
Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks
Mees, O., Hermann, L., Rosete-Beas, E., and Burgard, W · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Earlier work this paper cites.
Internvideo: General video foundation models via generative and discriminative learning
Wang, Y., Li, K., Li, Y., He, Y., Huang, B., Zhao, Z., Zhang, H., Xu, J., Liu, Y., Wang, Z., et al · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Stable video diffusion: Scaling latent video diffusion models to large datasets
Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D., Levi, Y., English, Z., Voleti, V., Letts, A., et al · 2023
Cited alongside, same era.
Videocrafter1: Open diffusion models for high-quality video generation, 2023
Chen, H., Xia, M., He, Y., Zhang, Y., Cun, X., Yang, S., Xing, J., Liu, Y., Chen, Q., Wang, X., Weng, C., and Shan, Y · 2023
Cited alongside, same era.
Dreamsim: Learning new dimensions of human visual similarity using synthetic data
Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Grauman, K., Westbury, A., Torresani, L., Kitani, K., Malik, J., Afouras, T., Ashutosh, K., Baiyya, V., Bansal, S., Boote, B., et al · 2024
Closest in time.
Seer: Language instructed video prediction with latent diffusion models, 2024
Gu, X., Wen, C., Ye, W., Song, J., and Gao, Y · 2024
Closest in time.
Llms meet multimodal generation and editing: A survey
He, Y., Liu, Z., Chen, J., Tian, Z., Liu, H., Chi, X., Liu, R., Yuan, R., Xing, Y., Wang, W., et al · 2024
Closest in time.
Minicpm: Unveiling the potential of small language models with scalable training strategies
Hu, S., Tu, Y., Han, X., He, C., Cui, G., Long, X., Zheng, Z., Fang, Y., Huang, Y., Zhao, W., et al · 2024
Closest in time.
Vbench: Comprehensive benchmark suite for video generative models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fu, S., Tamir, N., Sundaram, S., Chai, L., Zhang, R., Dekel, T., and Isola, P · 2023
Cited alongside, same era.
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Guo, Y., Yang, C., Rao, A., Liang, Z., Wang, Y., Qiao, Y., Agrawala, M., Lin, D., and Dai, B · 2023
Cited alongside, same era.
Learning to Act from Actionless Videos through Dense Correspondences
Ko, P.-C., Mao, J., Du, Y., Sun, S.-H., and Tenenbaum, J. B · 2023
Cited alongside, same era.
Videopoet: A large language model for zero-shot video generation
Kondratyuk, D., Yu, L., Gu, X., Lezama, J., Huang, J., Hornung, R., Adam, H., Akbari, H., Alon, Y., Birodkar, V., et al · 2023
Cited alongside, same era.
Mimic-it: Multi-modal in-context instruction tuning, 2023
Li, B., Zhang, Y., Chen, L., Wang, J., Pu, F., Yang, J., Li, C., and Liu, Z · 2023
Cited alongside, same era.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2023
Cited alongside, same era.
Open x-embodiment: Robotic learning datasets and rt-x models
Padalkar, A., Pooley, A., Jain, A., Bewley, A., Herzog, A., Irpan, A., Khazatsky, A., Rai, A., Singh, A., Brohan, A., et al · 2023
Cited alongside, same era.
Kosmos-2: Grounding multimodal large language models to the world
Peng, Z., Wang, W., Dong, L., Hao, Y., Huang, S., Ma, S., and Wei, F · 2023
Cited alongside, same era.
Huang, Z., He, Y., Yu, J., Zhang, F., Si, C., Jiang, Y., Zhang, Y., Wu, T., Jin, Q., Chanpaisit, N., et al · 2024
Closest in time.
Chat-univi: Unified visual representation empowers large language models with image and video understanding
Jin, P., Takanobu, R., Zhang, W., Cao, X., and Yuan, L · 2024
Closest in time.
Real-world robot applications of foundation models: A review, 2024
Kawaharazuka, K., Matsushima, T., Gambardella, A., Guo, J., Paxton, C., and Zeng, A · 2024
Closest in time.
Grounding video models to actions through goal conditioned exploration
Luo, Y. and Du, Y · 2024
Closest in time.
Video-chatgpt: Towards detailed video understanding via large vision and language models
Maaz, M., Rasheed, H., Khan, S., and Khan, F. S · 2024
Closest in time.
OpenAI · 2024
Closest in time.
Robovqa: Multimodal long-horizon reasoning for robotics
Sermanet, P., Ding, T., Zhao, J., Xia, F., Dwibedi, D., Gopalakrishnan, K., Chan, C., Dulac-Arnold, G., Maddineni, S., Joshi, N. J., et al · 2024
Closest in time.
Ego4d goal-step: Toward hierarchical understanding of procedural activities
Song, Y., Byrne, E., Nagarajan, T., Wang, H., Martin, M., and Torresani, L · 2024
Closest in time.
Aid: Adapting image2video diffusion models for instruction-guided video prediction
Xing, Z., Dai, Q., Weng, Z., Wu, Z., and Jiang, Y.-G · 2024
Closest in time.
Yang, A., Yang, B., Hui, B., Zheng, B., Yu, B., Zhou, C., Li, C., Li, C., Liu, D., Huang, F., et al · 2024
Closest in time.
Llava-next: A strong zero-shot video understanding model, April 2024c
Zhang, Y., Li, B., Liu, h., Lee, Y. j., Gui, L., Fu, D., Feng, J., Liu, Z., and Li, C · 2024
Closest in time.
Open-sora: Democratizing efficient video production for all, March 2024
Zheng, Z., Peng, X., Yang, T., Shen, C., Li, S., Liu, H., Zhou, Y., Li, T., and You, Y · 2024
Closest in time.
RoboDreamer: Learning compositional world models for robot imagination
Zhou, S., Du, Y., Chen, J., Li, Y., Yeung, D.-Y., and Gan, C · 2024
Closest in time.
Cosmos world foundation model platform for physical ai
Agarwal, N., Ali, A., Bala, M., Balaji, Y., Barker, E., Cai, T., Chattopadhyay, P., Chen, Y., Cui, Y., Ding, Y., et al · 2025
Closest in time.
Enerverse: Envisioning embodied future space for robotics manipulation
Huang, S., Chen, L., Zhou, P., Chen, S., Jiang, Z., Hu, Y., Gao, P., Li, H., Yao, M., and Ren, G · 2025
Closest in time.