Fetching the paper…
Reading the bibliography…
Synthetic videos nowadays is widely used to complement data scarcity and diversity of real-world videos.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
A survey on automatic image caption generation
Bai, S. and An, S · 2018
Earlier work this paper cites.
Show, recall, and tell: Image captioning with recall mechanism
Wang, L., Bai, Z., Zhang, Y., and Lu, H · 2020
Earlier work this paper cites.
A comprehensive study of deep video action recognition
Zhu, Y., Li, X., Liu, C., Zolfaghari, M., Xiong, Y., Wu, C., Zhang, Z., Tighe, J., Manmatha, R., and Li, M · 2020
Earlier work this paper cites.
Explain me the painting: Multi-topic knowledgeable art description generation
Bai, Z., Nakashima, Y., and Garcia, N · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Stable video diffusion: Scaling latent video diffusion models to large datasets
Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D., Levi, Y., English, Z., Voleti, V., Letts, A., et al · 2023
Earlier work this paper cites.
Unsupervised open-vocabulary object localization in videos
Fan, K., Bai, Z., Xiao, T., Zietlow, D., Horn, M., Zhao, Z., Simon-Gabriel, C.-J., Shou, M. Z., Locatello, F., Schiele, B., et al · 2023
Earlier work this paper cites.
Seed-bench: Benchmarking multimodal llms with generative comprehension
Li, B., Wang, R., Wang, G., Ge, Y., Ge, Y., and Shan, Y · 2023
Earlier work this paper cites.
Video-bench: A comprehensive benchmark and toolkit for evaluating video-based large language models
Ning, M., Zhu, B., Xie, Y., Lin, B., Cui, J., Yuan, L., Chen, D., and Yuan, L · 2023
Earlier work this paper cites.
Video understanding with large language models: A survey
Tang, Y., Bi, J., Xu, S., Song, L., Liang, S., Wang, T., Zhang, D., An, J., Lin, J., Zhu, R., et al · 2023
Earlier work this paper cites.
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
Wu, J. Z., Ge, Y., Wang, X., Lei, S. W., Gu, Y., Shi, Y., Hsu, W., Shan, Y., Qie, X., and Shou, M. Z · 2023
Earlier work this paper cites.
Videophy: Evaluating physical commonsense for video generation
Bansal, H., Lin, Z., Xie, T., Zong, Z., Yarom, M., Bitton, Y., Jiang, C., Sun, Y., Chang, K.-W., and Grover, A · 2024
Cited alongside, same era.
Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Fu, C., Dai, Y., Luo, Y., Li, L., Ren, S., Zhang, R., Wang, Z., Zhou, C., Shen, Y., Zhang, M., et al · 2024
Cited alongside, same era.
Ltx-video: Realtime video latent diffusion
HaCohen, Y., Chiprut, N., Brazowski, B., Shalem, D., Moshe, D., Richardson, E., Levin, E., Shiran, G., Zabari, N., Gordon, O., et al · 2024
Cited alongside, same era.
Vbench: Comprehensive benchmark suite for video generative models
Huang, Z., He, Y., Yu, J., Zhang, F., Si, C., Jiang, Y., Zhang, Y., Wu, T., Jin, Q., Chanpaisit, N., et al · 2024
Cited alongside, same era.
Janus: Decoupling visual encoding for unified multimodal understanding and generation
Wu, C., Chen, X., Wu, Z., Ma, Y., Liu, X., Pan, Z., Liu, W., Xie, Z., Yu, X., Ruan, C., et al · 2024
Later among the works it cites.
Show-o: One single transformer to unify multimodal understanding and generation
Xie, J., Mao, W., Bai, Z., Zhang, D. J., Wang, W., Lin, K. Q., Gu, Y., Chen, Z., Yang, Z., and Shou, M. Z · 2024
Later among the works it cites.
Loki: A comprehensive synthetic data detection benchmark using large multimodal models
Ye, J., Zhou, B., Huang, Z., Zhang, J., Bai, T., Kang, H., He, J., Lin, H., Wang, Z., Wu, T., et al · 2024
Later among the works it cites.
The dawn of video generation: Preliminary explorations with sora-like models
Zeng, A., Yang, Y., Chen, W., and Liu, W · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hurst, A., Lerer, A., Goucher, A. P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A., Welihinda, A., Hayes, A., Radford, A., et al · 2024
Cited alongside, same era.
Pyramidal flow matching for efficient video generative modeling
Jin, Y., Sun, Z., Li, N., Xu, K., Jiang, H., Zhuang, N., Huang, Q., Song, Y., Mu, Y., and Lin, Z · 2024
Cited alongside, same era.
Hunyuanvideo: A systematic framework for large video generative models
Kong, W., Tian, Q., Zhang, Z., Min, R., Dai, Z., Zhou, J., Xiong, J., Li, X., Wu, B., Zhang, J., et al · 2024
Cited alongside, same era.
Mvbench: A comprehensive multi-modal video understanding benchmark
Li, K., Wang, Y., He, Y., Li, Y., Wang, Y., Liu, Y., Wang, Z., Xu, J., Chen, G., Luo, P., et al · 2024
Cited alongside, same era.
Openvid-1m: A large-scale high-quality dataset for text-to-video generation
Nan, K., Xie, R., Zhou, P., Fan, T., Yang, Z., Chen, Z., Li, X., Yang, J., and Tai, Y · 2024
Cited alongside, same era.
Longvu: Spatiotemporal adaptive compression for long video-language understanding
Shen, X., Xiong, Y., Zhao, C., Wu, L., Chen, J., Zhu, C., Liu, Z., Xiao, F., Varadarajan, B., Bordes, F., et al · 2024
Cited alongside, same era.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Team, G., Georgiev, P., Lei, V. I., Burnell, R., Bai, L., Gulati, A., Tanzer, G., Vincent, D., Pan, Z., Wang, S., et al · 2024
Cited alongside, same era.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Wang, P., Bai, S., Tan, S., Wang, S., Fan, Z., Bai, J., Chen, K., Liu, X., Wang, J., Ge, W., et al · 2024
Cited alongside, same era.
Show-1: Marrying pixel and latent diffusion models for text-to-video generation
Zhang, D. J., Wu, J. Z., Liu, J.-W., Zhao, R., Ran, L., Gu, Y., Gao, D., and Shou, M. Z · 2024
Later among the works it cites.
Cosmos world foundation model platform for physical ai
Agarwal, N., Ali, A., Bala, M., Balaji, Y., Barker, E., Cai, T., Chattopadhyay, P., Chen, Y., Cui, Y., Ding, Y., et al · 2025
Closest in time.
One token to seg them all: Language instructed reasoning segmentation in videos
Bai, Z., He, T., Mei, H., Wang, P., Gao, Z., Chen, J., Zhang, Z., and Shou, M. Z · 2025
Closest in time.
Mochi 1: A new sota in open-source video generation models
GenmoAI · 2025
Closest in time.
Hailuo ai: Transform idea to visual with ai
Hailuo · 2025
Closest in time.
Kling ai: Next-generation ai creative studio
KLING · 2025
Closest in time.
Luma dream machine
LumaLabs · 2025
Closest in time.
Genvidbench: A challenging benchmark for detecting ai-generated video
Ni, Z., Yan, Q., Huang, M., Yuan, T., Tang, Y., Hu, H., Chen, X., and Wang, Y · 2025
Closest in time.