Fetching the paper…
Reading the bibliography…
We present the ShareGPT4Video series, aiming to facilitate the video understanding of large video-language models (LVLMs) and the video generation of text-to-video models (T2VMs) via dense and precise captions.
Adaptive key frame extraction using unsupervised clustering
Y. Zhuang, Y. Rui, T. S. Huang, and S. Mehrotra · 1998
Earlier work this paper cites.
Efficient key-frame extraction and video analysis
J. Calic and E. Izuierdo · 2002
Earlier work this paper cites.
Tgif: A new dataset and benchmark on animated gif description
Y. Li, Y. Song, L. Cao, J. Tetreault, L. Goldberg, A. Jaimes, and J. Luo · 2016
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
D. Xu, Z. Zhao, J. Xiao, F. Wu, H. Zhang, X. He, and Y. Zhuang · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Clevrer: Collision events for video representation and reasoning
K. Yi, C. Gan, Y. Li, P. Kohli, J. Wu, A. Torralba, and J. B. Tenenbaum · 2019
Earlier work this paper cites.
Activitynet-qa: A dataset for understanding complex web videos via question answering
Z. Yu, D. Xu, J. Yu, T. Yu, Z. Zhao, Y. Zhuang, and D. Tao · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
Bdd100k: A diverse driving dataset for heterogeneous multitask learning
F. Yu, H. Chen, X. Wang, W. Xian, Y. Chen, F. Liu, V. Madhavan, and T. Darrell · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Next-qa: Next phase of question-answering to explaining temporal actions
J. Xiao, X. Shang, A. Yao, and T.-S. Chua · 2021
Earlier work this paper cites.
Ego4d: Around the world in 3,000 hours of egocentric video
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, et al · 2022
Earlier work this paper cites.
Imagen video: High definition video generation with diffusion models
J. Ho, W. Chan, C. Saharia, J. Whang, R. Gao, A. Gritsenko, D. P. Kingma, B. Poole, M. Norouzi, D. J. Fleet, et al · 2022
Earlier work this paper cites.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
W. Hong, M. Ding, W. Zheng, X. Liu, and J. Tang · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Earlier work this paper cites.
Make-a-video: Text-to-video generation without text-video data
U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni, et al · 2022
Earlier work this paper cites.
Magicvideo: Efficient video generation with latent diffusion models
D. Zhou, W. Wang, H. Yan, W. Lv, Y. Zhu, and J. Feng · 2022
Earlier work this paper cites.
Qwen-vl: A frontier large vision-language model with versatile abilities
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou · 2023
Earlier work this paper cites.
Improving image generation with better captions
J. Betker, G. Goh, L. Jing, T. Brooks, J. Wang, L. Li, L. Ouyang, J. Zhuang, J. Lee, Y. Guo, et al · 2023
Earlier work this paper cites.
Stable video diffusion: Scaling latent video diffusion models to large datasets
A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, D. Lorenz, Y. Levi, Z. English, V. Voleti, A. Letts, et al · 2023
Earlier work this paper cites.
Muse: Text-to-image generation via masked generative transformers
H. Chang, H. Zhang, J. Barber, A. Maschinot, J. Lezama, L. Jiang, M.-H. Yang, K. Murphy, W. T. Freeman, M. Rubinstein, et al · 2023
Earlier work this paper cites.
Pixart-alpha: Fast training of diffusion transformer for photorealistic text-to-image synthesis
J. Chen, J. Yu, C. Ge, L. Yao, E. Xie, Y. Wu, Z. Wang, J. Kwok, P. Luo, H. Lu, et al · 2023
Earlier work this paper cites.
Video chatcaptioner: Towards the enriched spatiotemporal descriptions
J. Chen, D. Zhu, K. Haydarov, X. Li, and M. Elhoseiny · 2023
Earlier work this paper cites.
Sharegpt4v: Improving large multi-modal models with better captions
L. Chen, J. Li, X. Dong, P. Zhang, C. He, J. Wang, F. Zhao, and D. Lin · 2023
Cited alongside, same era.
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, Z. Muyan, Q. Zhang, X. Zhu, L. Lu, et al · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, et al · 2023
Cited alongside, same era.
Improving image generation with better captions
B. James, G. Gabriel, J. Li, and B. Tim · 2023
Cited alongside, same era.
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
J. Z. Wu, Y. Ge, X. Wang, S. W. Lei, Y. Gu, Y. Shi, W. Hsu, Y. Shan, X. Qie, and M. Z. Shou · 2023
Later among the works it cites.
mplug-owl: Modularization empowers large language models with multimodality
Q. Ye, H. Xu, G. Xu, J. Ye, M. Yan, Y. Zhou, J. Wang, A. Hu, P. Shi, Y. Shi, et al · 2023
Later among the works it cites.
Make pixels dance: High-dynamic video generation
Y. Zeng, G. Wei, J. Zheng, J. Zou, Y. Wei, Y. Zhang, and H. Li · 2023
Later among the works it cites.
Video-llama: An instruction-tuned audio-visual language model for video understanding
H. Zhang, X. Li, and L. Bing · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
L. Zhang, A. Rao, and M. Agrawala · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Jin, R. Takanobu, C. Zhang, X. Cao, and L. Yuan · 2023
Cited alongside, same era.
Otter: A multi-modal model with in-context instruction tuning
B. Li, Y. Zhang, L. Chen, J. Wang, J. Yang, and Z. Liu · 2023
Cited alongside, same era.
J. Li, D. Li, S. Savarese, and S. Hoi · 2023
Cited alongside, same era.
Videochat: Chat-centric video understanding
K. Li, Y. He, Y. Wang, Y. Li, W. Wang, P. Luo, Y. Wang, L. Wang, and Y. Qiao · 2023
Cited alongside, same era.
Mvbench: A comprehensive multi-modal video understanding benchmark
K. Li, Y. Wang, Y. He, Y. Li, Y. Wang, Y. Liu, Z. Wang, J. Xu, G. Chen, P. Luo, et al · 2023
Cited alongside, same era.
Llama-vid: An image is worth 2 tokens in large language models
Y. Li, C. Wang, and J. Jia · 2023
Cited alongside, same era.
Monkey: Image resolution and text label are important things for large multi-modal models
Z. Li, B. Yang, Q. Liu, Z. Ma, S. Zhang, J. Yang, Y. Sun, Y. Liu, and X. Bai · 2023
Cited alongside, same era.
Video-llava: Learning united visual representation by alignment before projection
B. Lin, B. Zhu, Y. Ye, M. Ning, P. Jin, and L. Yuan · 2023
Cited alongside, same era.
Later among the works it cites.
Llama-adapter: Efficient fine-tuning of language models with zero-init attention
R. Zhang, J. Han, A. Zhou, X. Hu, S. Yan, P. Lu, H. Li, P. Gao, and Y. Qiao · 2023
Later among the works it cites.
K. Ataallah, X. Shen, E. Abdelrahman, E. Sleiman, D. Zhu, J. Ding, and M. Elhoseiny · 2024
Closest in time.
Allava: Harnessing gpt4v-synthesized data for a lite vision-language model
G. H. Chen, S. Chen, R. Zhang, J. Chen, X. Wu, Z. Zhang, Z. Chen, J. Li, X. Wan, and B. Wang · 2024
Closest in time.
Are we on the right way for evaluating large vision-language models?
L. Chen, J. Li, X. Dong, P. Zhang, Y. Zang, Z. Chen, H. Duan, J. Wang, Y. Qiao, D. Lin, et al · 2024
Closest in time.
Panda-70m: Captioning 70m videos with multiple cross-modality teachers
T.-S. Chen, A. Siarohin, W. Menapace, E. Deyneka, H.-w. Chao, B. E. Jeon, Y. Fang, H.-Y. Lee, J. Ren, M.-H. Yang, et al · 2024
Closest in time.
X. Dong, P. Zhang, Y. Zang, Y. Cao, B. Wang, L. Ouyang, X. Wei, S. Zhang, H. Duan, M. Cao, et al · 2024
Closest in time.
X. Dong, P. Zhang, Y. Zang, Y. Cao, B. Wang, L. Ouyang, S. Zhang, H. Duan, W. Zhang, Y. Li, et al · 2024
Closest in time.
An image grid can be worth a video: Zero-shot video question answering using a vlm
W. Kim, C. Choi, W. Lee, and W. Rhee · 2024
Closest in time.
Open-sora-plan, Apr. 2024
P.-Y. Lab and T. A. etc · 2024
Closest in time.
Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024
B. Li, K. Zhang, H. Zhang, D. Guo, R. Zhang, F. Li, Y. Zhang, Z. Liu, and C. Li · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee · 2024
Closest in time.
Tempcompass: Do video llms really understand videos?
Y. Liu, S. Li, Y. Liu, Y. Wang, S. Ren, L. Li, S. Chen, X. Sun, and L. Hou · 2024
Closest in time.
Sora: A review on background, technology, limitations, and opportunities of large vision models
Y. Liu, K. Zhang, Y. Li, Z. Yan, C. Gao, R. Chen, Z. Yuan, Y. Huang, H. Sun, J. Gao, et al · 2024
Closest in time.
Latte: Latent diffusion transformer for video generation
X. Ma, Y. Wang, G. Jia, X. Chen, Z. Liu, Y.-F. Li, C. Chen, and Y. Qiao · 2024
Closest in time.
Q-Bench: A benchmark for general-purpose foundation models on low-level vision
H. Wu, Z. Zhang, E. Zhang, C. Chen, L. Liao, A. Wang, C. Li, W. Sun, Q. Yan, G. Zhai, and W. Lin · 2024
Closest in time.
Q-Instruct: Improving low-level visual abilities for multi-modality foundation models
H. Wu, Z. Zhang, E. Zhang, C. Chen, L. Liao, A. Wang, K. Xu, C. Li, J. Hou, G. Zhai, et al · 2024
Closest in time.
Yi: Open foundation models by 01. ai
A. Young, B. Chen, C. Li, C. Huang, G. Zhang, G. Zhang, H. Li, J. Zhu, J. Chen, J. Chang, et al · 2024
Closest in time.
Tinyllava: A framework of small-scale large multimodal models
B. Zhou, Y. Hu, X. Weng, J. Jia, J. Luo, X. Liu, J. Wu, and L. Huang · 2024
Closest in time.