Fetching the paper…
Reading the bibliography…
In light of recent advances in multimodal Large Language Models (LLMs), there is increasing attention to scaling them from image-text data to more informative real-world videos.
MPEG: A video compression standard for multimedia applications
Le Gall, D · 1991
Earlier work this paper cites.
The computation of optical flow
Beauchemin, S. S. and Barron, J. L · 1995
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
Chen, D. and Dolan, W. B · 2011
Earlier work this paper cites.
Im2Text: Describing images using 1 million captioned photographs
Ordonez, V., Kulkarni, G., and Berg, T. L · 2011
Earlier work this paper cites.
A dataset of 101 human action classes from videos in the wild
Soomro, K., Zamir, A. R., and Shah, M · 2012
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3D convolutional networks
Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M · 2015
Earlier work this paper cites.
MSR-VTT: A large video description dataset for bridging video and language
Xu, J., Mei, T., Yao, T., and Rui, Y · 2016
Earlier work this paper cites.
Real-time action recognition with enhanced motion vector CNNs
Zhang, B., Wang, L., Wang, Z., Qiao, Y., and Wang, H · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J. and Zisserman, A · 2017
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
GANs trained by a two time-scale update rule converge to a local Nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
The Kinetics human action video dataset
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., et al · 2017
Earlier work this paper cites.
Neural discrete representation learning
van den Oord, A., Vinyals, O., and Kavukcuoglu, K · 2017
Earlier work this paper cites.
VizWiz grand challenge: Answering visual questions from blind people
Gurari, D., Li, Q., Stangl, A. J., Guo, A., Lin, C., Grauman, K., Luo, J., and Bigham, J. P · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Unterthiner, T., Van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S · 2018
Earlier work this paper cites.
Compressed video action recognition
Wu, C.-Y., Zaheer, M., Hu, H., Manmatha, R., Smola, A. J., and Krähenbühl, P · 2018
Earlier work this paper cites.
Character region awareness for text detection
Baek, Y., Lee, B., Han, D., Yun, S., and Lee, H · 2019
Earlier work this paper cites.
GQA: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A. and Manning, C. D · 2019
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Song, Y. and Ermon, S · 2019
Earlier work this paper cites.
ActivityNet-QA: A dataset for understanding complex web videos via question answering
Yu, Z., Xu, D., Yu, J., Yu, T., Zhao, Z., Zhuang, Y., and Tao, D · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Train sparsely, generate densely: Memory-efficient unsupervised training of high-resolution temporal GAN
Saito, M., Saito, S., Koyama, M., and Kobayashi, S · 2020
Cited alongside, same era.
Denoising diffusion implicit models
Song, J., Meng, C., and Ermon, S · 2020
Cited alongside, same era.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Bain, M., Nagrani, A., Varol, G., and Zisserman, A · 2021
Cited alongside, same era.
Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Changpinyo, S., Sharma, P., Ding, N., and Soricut, R · 2021
Cited alongside, same era.
Taming transformers for high-resolution image synthesis
Esser, P., Rombach, R., and Ommer, B · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Preserve your own correlation: A noise prior for video diffusion models
Ge, S., Nah, S., Liu, G., Poon, T., Tao, A., Catanzaro, B., Jacobs, D., Huang, J.-B., Liu, M.-Y., and Balaji, Y · 2023
Later among the works it cites.
Gemini: A family of highly capable multimodal models
Gemini Team, G · 2023
Later among the works it cites.
Emu video: Factorizing text-to-video generation by explicit image conditioning
Girdhar, R., Singh, M., Brown, A., Duval, Q., Azadi, S., Rambhatla, S. S., Shah, A., Yin, X., Parikh, D., and Misra, I · 2023
Later among the works it cites.
CogVideo: Large-scale pretraining for text-to-video generation via transformers
Hong, W., Ding, M., Zheng, W., Liu, X., and Tang, J · 2023
Later among the works it cites.
VideoPoet: A large language model for zero-shot video generation
Kondratyuk, D., Yu, L., Gu, X., Lezama, J., Huang, J., Hornung, R., Adam, H., Akbari, H., Alon, Y., Birodkar, V., et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
GODIVA: Generating open-domain videos from natural descriptions
Wu, C., Huang, L., Zhang, Q., Li, B., Ji, L., Yang, F., Sapiro, G., and Duan, N · 2021
Cited alongside, same era.
VideoGPT: Video generation using VQ-VAE and transformers
Yan, W., Zhang, Y., Abbeel, P., and Srinivas, A · 2021
Cited alongside, same era.
CM3: A causal masked multimodal model of the internet
Aghajanyan, A., Huang, B., Ross, C., Karpukhin, V., Xu, H., Goyal, N., Okhonko, D., Joshi, M., Ghosh, G., Lewis, M., et al · 2022
Cited alongside, same era.
Flamingo: A visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Cited alongside, same era.
Video diffusion models
Ho, J., Salimans, T., Gritsenko, A., Chan, W., Norouzi, M., and Fleet, D. J · 2022
Cited alongside, same era.
Later among the works it cites.
The role of ImageNet classes in Fréchet Inception distance
Kynkäänniemi, T., Karras, T., Aittala, M., Aila, T., and Lehtinen, J · 2023
Later among the works it cites.
Lian, L., Li, B., Yala, A., and Darrell, T · 2023
Later among the works it cites.
Video-LLaVA: Learning united visual representation by alignment before projection
Lin, B., Zhu, B., Ye, Y., Ning, M., Jin, P., and Yuan, L · 2023
Later among the works it cites.
Video-ChatGPT: Towards detailed video understanding via large vision and language models
Maaz, M., Rasheed, H., Khan, S., and Khan, F. S · 2023
Later among the works it cites.
Freenoise: Tuning-free longer video diffusion via noise rescheduling
Qiu, H., Xia, M., Zhang, Y., He, Y., Wang, X., Shan, Y., and Liu, Z · 2023
Later among the works it cites.
Make-A-Video: Text-to-video generation without text-video data
Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., et al · 2023
Later among the works it cites.
RedPajama: an open dataset for training large language models, 2023
Together Computer · 2023
Later among the works it cites.
Phenaki: Variable length video generation from open domain textual descriptions
Villegas, R., Babaeizadeh, M., Kindermans, P.-J., Moraldo, H., Zhang, H., Saffar, M. T., Castro, S., Kunze, J., and Erhan, D · 2023
Later among the works it cites.
mPLUG-Owl: Modularization empowers large language models with multimodality
Ye, Q., Xu, H., Xu, G., Ye, J., Yan, M., Zhou, Y., Wang, J., Hu, A., Shi, P., Shi, Y., et al · 2023
Later among the works it cites.
Make pixels dance: High-dynamic video generation
Zeng, Y., Wei, G., Zheng, J., Zou, J., Wei, Y., Zhang, Y., and Li, H · 2023
Later among the works it cites.
DreamLLM: Synergistic multimodal comprehension and creation
Dong, R., Han, C., Peng, Y., Qi, Z., Ge, Z., Yang, J., Zhao, L., Sun, J., Zhou, H., Wei, H., et al · 2024
Closest in time.
LayoutGPT: Compositional visual planning and generation with large language models
Feng, W., Zhu, W., Fu, T.-j., Jampani, V., Akula, A., He, X., Basu, S., Wang, X. E., and Wang, W. Y · 2024
Closest in time.
Unified language-vision pretraining in LLM with dynamic discrete visual tokenization
Jin, Y., Xu, K., Xu, K., Chen, L., Liao, C., Tan, J., Huang, Q., Chen, B., Lei, C., Liu, A., et al · 2024
Closest in time.
Egoschema: A diagnostic benchmark for very long-form video language understanding
Mangalam, K., Akshulakov, R., and Malik, J · 2024
Closest in time.
Perception test: A diagnostic benchmark for multimodal video models
Patraucean, V., Smaira, L., Gupta, A., Recasens, A., Markeeva, L., Banarse, D., Koppula, S., Malinowski, M., Yang, Y., Doersch, C., et al · 2024
Closest in time.
SDXL: Improving latent diffusion models for high-resolution image synthesis
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R · 2024
Closest in time.
Decouple content and motion for conditional image-to-video generation
Shen, C., Gan, Y., Chen, C., Zhu, X., Cheng, L., and Wang, J · 2024
Closest in time.
Emu: Generative pretraining in multimodality
Sun, Q., Yu, Q., Cui, Y., Zhang, F., Zhang, X., Wang, Y., Gao, H., Liu, J., Huang, T., and Wang, X · 2024
Closest in time.
InternVid: A large-scale video-text dataset for multimodal understanding and generation
Wang, Y., He, Y., Li, Y., Li, K., Yu, J., Ma, X., Chen, X., Wang, Y., Luo, P., Liu, Z., et al · 2024
Closest in time.
Video-LLaMA: An instruction-tuned audio-visual language model for video understanding
Zhang, H., Li, X., and Bing, L · 2024
Closest in time.