Fetching the paper…
Reading the bibliography…
Long-form video processing fundamentally challenges vision-language models (VLMs) due to the high computational costs of handling extended temporal sequences.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Soomro, K., Zamir, A. R., and Shah, M · 2012
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Caba Heilbron, F., Escorcia, V., Ghanem, B., and Carlos Niebles, J · 2015
Earlier work this paper cites.
The” something something” video database for learning and evaluating visual common sense
Goyal, R., Ebrahimi Kahou, S., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., Mueller-Freitag, M., et al · 2017
Earlier work this paper cites.
Mixed precision training
Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D., Ginsburg, B., Houston, M., Kuchaiev, O., Venkatesh, G., and Wu, H · 2018
Earlier work this paper cites.
Adaframe: Adaptive frame selection for fast video recognition
Wu, Z., Xiong, C., Ma, C.-Y., Socher, R., and Davis, L. S · 2019
Earlier work this paper cites.
Clevrer: Collision events for video representation and reasoning
Yi, K., Gan, C., Li, Y., Kohli, P., Wu, J., Torralba, A., and Tenenbaum, J. B · 2019
Earlier work this paper cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Bain, M., Nagrani, A., Varol, G., and Zisserman, A · 2021
Earlier work this paper cites.
Detecting moments and highlights in videos via natural language queries
Lei, J., Berg, T. L., and Bansal, M · 2021
Earlier work this paper cites.
A benchmark for situated reasoning in real-world videos
Wu, B. and Star, S. Y · 2021
Earlier work this paper cites.
Next-qa: Next phase of question-answering to explaining temporal actions
Xiao, J., Shang, X., Yao, A., and Chua, T.-S · 2021
Earlier work this paper cites.
Mgsampler: An explainable sampling strategy for video action recognition
Zhi, Y., Tong, Z., Wang, L., and Wu, G · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Earlier work this paper cites.
Video pretraining (VPT): Learning to act by watching unlabeled online videos
Baker, B., Akkaya, I., Zhokov, P., Huizinga, J., Tang, J., Ecoffet, A., Houghton, B., Sampedro, R., and Clune, J · 2022
Earlier work this paper cites.
Egotaskqa: Understanding human tasks in egocentric videos
Jia, B., Lei, T., Zhu, S.-C., and Huang, S · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J · 2023
Earlier work this paper cites.
Token merging: Your vit but faster
Bolya, D., Fu, C.-Y., Dai, X., Zhang, P., Feichtenhofer, C., and Hoffman, J · 2023
Earlier work this paper cites.
Videochat: Chat-centric video understanding
KunChang, L., Yinan, H., Yi, W., Yizhuo, L., Wenhai, W., Ping, L., Yali, W., Limin, W., and Yu, Q · 2023
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Li, J., Li, D., Savarese, S., and Hoi, S · 2023
Cited alongside, same era.
Video-llava: Learning united visual representation by alignment before projection
Lin, B., Ye, Y., Zhu, B., Cui, J., Ning, M., Jin, P., and Yuan, L · 2023
Cited alongside, same era.
Egoschema: A diagnostic benchmark for very long-form video language understanding
Mangalam, K., Akshulakov, R., and Malik, J · 2023
Cited alongside, same era.
Perception test: A diagnostic benchmark for multimodal video models
Patraucean, V., Smaira, L., Gupta, A., Recasens, A., Markeeva, L., Banarse, D., Koppula, S., Malinowski, M., Yang, Y., Doersch, C., et al · 2023
Fei, J., Li, D., Deng, Z., Wang, Z., Liu, G., and Wang, H · 2024
Later among the works it cites.
Chat-univi: Unified visual representation empowers large language models with image and video understanding
Jin, P., Takanobu, R., Zhang, W., Cao, X., and Yuan, L · 2024
Later among the works it cites.
In search of needles in a 11m haystack: Recurrent memory finds what llms miss, 2024
Kuratov, Y., Bulatov, A., Anokhin, P., Sorokin, D., Sorokin, A., and Burtsev, M · 2024
Later among the works it cites.
Vila: On pre-training for visual language models
Lin, J., Yin, H., Ping, W., Molchanov, P., Shoeybi, M., and Han, S · 2024
Later among the works it cites.
Openvid-1m: A large-scale high-quality dataset for text-to-video generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Moviechat: From dense token to sparse memory for long video understanding
Song, E., Chai, W., Wang, G., Zhang, Y., Zhou, H., Wu, F., Guo, X., Ye, T., Lu, Y., Hwang, J.-N., et al · 2023
Cited alongside, same era.
Internvid: A large-scale video-text dataset for multimodal understanding and generation
Wang, Y., He, Y., Li, Y., Li, K., Yu, J., Ma, X., Li, X., Chen, G., Chen, X., Wang, Y., et al · 2023
Cited alongside, same era.
Daydreamer: World models for physical robot learning
Wu, P., Escontrela, A., Hafner, D., Abbeel, P., and Goldberg, K · 2023
Cited alongside, same era.
mplug-owl: Modularization empowers large language models with multimodality
Ye, Q., Xu, H., Xu, G., Ye, J., Yan, M., Zhou, Y., Wang, J., Hu, A., Shi, P., Shi, Y., et al · 2023
Cited alongside, same era.
Hierarchical video-moment retrieval and step-captioning
Zala, A., Cho, J., Kottur, S., Chen, X., Oguz, B., Mehdad, Y., and Bansal, M · 2023
Cited alongside, same era.
Sigmoid loss for language image pre-training
Zhai, X., Mustafa, B., Kolesnikov, A., and Beyer, L · 2023
Cited alongside, same era.
Video-llama: An instruction-tuned audio-visual language model for video understanding
Zhang, H., Li, X., and Bing, L · 2023
Cited alongside, same era.
Nan, K., Xie, R., Zhou, P., Fan, T., Yang, Z., Chen, Z., Li, X., Yang, J., and Tai, Y · 2024
Later among the works it cites.
Hello gpt-4o
OpenAI · 2024
Later among the works it cites.
Cinepile: A long video question answering dataset and benchmark
Rawal, R., Saifullah, K., Farré, M., Basri, R., Jacobs, D., Somepalli, G., and Goldstein, T · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T., Alayrac, J.-b., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., et al · 2024
Later among the works it cites.
Longvu: Spatiotemporal adaptive compression for long video-language understanding
Shen, X., Xiong, Y., Zhao, C., Wu, L., Chen, J., Zhu, C., Liu, Z., Xiao, F., Varadarajan, B., Bordes, F., et al · 2024
Later among the works it cites.
Video-xl: Extra-long vision language model for hour-scale video understanding
Shu, Y., Zhang, P., Liu, Z., Qin, M., Zhou, J., Huang, T., and Zhao, B · 2024
Later among the works it cites.
Moviechat: From dense token to sparse memory for long video understanding
Song, E., Chai, W., Wang, G., Zhang, Y., Zhou, H., Wu, F., Chi, H., Guo, X., Ye, T., Zhang, Y., et al · 2024
Later among the works it cites.
Longvideobench: A benchmark for long-context interleaved video-language understanding, 2024
Wu, H., Li, D., Chen, B., and Li, J · 2024
Later among the works it cites.
Pllava: Parameter-free llava extension from images to videos for video dense captioning
Xu, L., Zhao, Y., Zhou, D., Lin, Z., Ng, S. K., and Feng, J · 2024
Later among the works it cites.
Minicpm-v: A gpt-4v level mllm on your phone
Yao, Y., Yu, T., Zhang, A., Wang, C., Cui, J., Zhu, H., Cai, T., Li, H., Zhao, W., He, Z., et al · 2024
Later among the works it cites.
mplug-owl3: Towards long image-sequence understanding in multi-modal large language models
Ye, J., Xu, H., Liu, H., Hu, A., Yan, M., Qian, Q., Zhang, J., Huang, F., and Zhou, J · 2024
Later among the works it cites.
Llama-vid: An image is worth 2 tokens in large language models
Li, Y., Wang, C., and Jia, J · 2025
Closest in time.
Llava-mini: Efficient image and video large multimodal models with one vision token, 2025
Zhang, S., Fang, Q., Yang, Z., and Feng, Y · 2025
Closest in time.