Fetching the paper…
Reading the bibliography…
Multimodal large language models (MLLMs) demonstrate remarkable capabilities in handling complex multimodal tasks and are increasingly adopted in video understanding applications.
A mathematical theory of communication
Claude E Shannon · 1948
Earlier work this paper cites.
On measures of entropy and information
Alfréd Rényi · 1961
Earlier work this paper cites.
New non-additive measures of entropy for discrete probability distributions
Bhudev D Sharma and Dharam P Mittal · 1975
Earlier work this paper cites.
Possible generalization of boltzmann-gibbs statistics
Constantino Tsallis · 1988
Earlier work this paper cites.
On generalized information measures and their applications
Inder Jeet Taneja · 1989
Earlier work this paper cites.
A summary on entropy statistics
Maria Dolores Esteban and Domingo Morales · 1995
Earlier work this paper cites.
Polynomial expansion for orientation and motion estimation
Gunnar Farnebäck · 2002
Earlier work this paper cites.
Generalised information and entropy measures in physics
Christian Beck · 2009
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
Logan: Membership inference attacks against generative models
Jamie Hayes, Luca Melis, George Danezis, and Emiliano De Cristofaro · 2017
Earlier work this paper cites.
Learning latent subevents in activity videos using temporal attention filters
A Piergiovanni, Chenyou Fan, and Michael Ryoo · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov · 2017
Earlier work this paper cites.
Machine learning models that remember too much
Congzheng Song, Thomas Ristenpart, and Vitaly Shmatikov · 2017
Earlier work this paper cites.
Ahmed Salem, Yang Zhang, Mathias Humbert, Pascal Berrang, Mario Fritz, and Michael Backes · 2018
Earlier work this paper cites.
Privacy risk in machine learning: Analyzing the connection to overfitting
Samuel Yeom, Irene Giacomelli, Matt Fredrikson, and Somesh Jha · 2018
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song · 2019
Earlier work this paper cites.
Membership inference attacks against adversarially robust deep learning models
Liwei Song, Reza Shokri, and Prateek Mittal · 2019
Earlier work this paper cites.
Gan-leaks: A taxonomy of membership inference attacks against generative models
Dingfan Chen, Ning Yu, Yang Zhang, and Mario Fritz · 2020
Earlier work this paper cites.
Information leakage in embedding models
Congzheng Song and Ananth Raghunathan · 2020
Earlier work this paper cites.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Earlier work this paper cites.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut · 2021
Earlier work this paper cites.
A hierarchical variational neural uncertainty model for stochastic video prediction
Moitreya Chatterjee, Narendra Ahuja, and Anoop Cherian · 2021
Cited alongside, same era.
Next-qa: Next phase of question answering to explaining temporal actions
Xiao Lin and Chenliang Xu · 2021
Cited alongside, same era.
Membership inference on word embedding and beyond
Saeed Mahloujifar, Huseyin A Inan, Melissa Chase, Esha Ghosh, and Marcello Hasegawa · 2021
Cited alongside, same era.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki · 2021
Cited alongside, same era.
Systematic evaluation of privacy risks of machine learning models
Liwei Song and Prateek Mittal · 2021
Cited alongside, same era.
Vision-language models for medical report generation and visual question answering: A review
Iryna Hartsock and Ghulam Rasool · 2024
Later among the works it cites.
Bliva: A simple multimodal llm for better handling of text-rich visual questions
Wenbo Hu, Yifan Xu, Yi Li, Weiyue Li, Zeyuan Chen, and Zhuowen Tu · 2024
Later among the works it cites.
Data lineage inference: Uncovering privacy vulnerabilities of dataset pruning
Qi Li, Cheng-Long Wang, Yinzhi Cao, and Di Wang · 2024
Later among the works it cites.
Membership inference attacks against large vision-language models
Zhan Li, Yongtao Wu, Yihang Chen, Francesco Tonin, Elias Abad Rocamora, and Volkan Cevher · 2024
Later among the works it cites.
Improved baselines with visual instruction tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Cited alongside, same era.
Membership inference attacks from first principles
Nicholas Carlini, Steve Chien, Milad Nasr, Shuang Song, Andreas Terzis, and Florian Tramer · 2022
Cited alongside, same era.
Expanding language-image pretrained models for general video recognition
Bolin Ni, Houwen Peng, Minghao Chen, Songyang Zhang, Gaofeng Meng, Jianlong Fu, Shiming Xiang, and Haibin Ling · 2022
Cited alongside, same era.
Optimizing video prediction via video frame interpolation
Yue Wu, Qiang Wen, and Qifeng Chen · 2022
Cited alongside, same era.
A survey on deep learning technique for video segmentation
Tianfei Zhou, Fatih Porikli, David J Crandall, Luc Van Gool, and Wenguan Wang · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Ziqing Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al · 2023
Cited alongside, same era.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee · 2024
Later among the works it cites.
St-llm: Large language models are effective temporal learners
Ruyang Liu, Chen Li, Haoran Tang, Yixiao Ge, Ying Shan, and Ge Li · 2024
Later among the works it cites.
Tempcompass: Do video llms really understand videos?
Yuanxin Liu, Shicheng Li, Yi Liu, Yuxiang Wang, Shuhuai Ren, Lei Li, Sishuo Chen, Xu Sun, and Lu Hou · 2024
Later among the works it cites.
Slowfocus: Enhancing fine-grained temporal understanding in video llm
Ming Nie, Dan Ding, Chunwei Wang, Yuanfan Guo, Jianhua Han, Hang Xu, and Li Zhang · 2024
Later among the works it cites.
Cinepile: A long video question answering dataset and benchmark
Ruchit Rawal, Khalid Saifullah, Ronen Basri, David Jacobs, Gowthami Somepalli, and Tom Goldstein · 2024
Later among the works it cites.
Video-xl: Extra-long vision language model for hour-scale video understanding
Yan Shu, Peitian Zhang, Zheng Liu, Minghao Qin, Junjie Zhou, Tiejun Huang, and Bo Zhao · 2024
Later among the works it cites.
Cheng-Long Wang, Qi Li, Zihang Xiang, Yinzhi Cao, and Di Wang · 2024
Later among the works it cites.
Cogvlm: Visual expert for pretrained language models
Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Song XiXuan, et al · 2024
Later among the works it cites.
Videoagent: Long-form video understanding with large language model as agent
Xiaohan Wang, Yuhui Zhang, Orr Zohar, and Serena Yeung-Levy · 2024
Later among the works it cites.
Min-k%++: Improved baseline for detecting pre-training data from large language models
Jingyang Zhang, Jingwei Sun, Eric Yeats, Yang Ouyang, Martin Kuo, Jianyi Zhang, Hao Frank Yang, and Hai Li · 2024
Later among the works it cites.
Long context transfer from language to vision
Peiyuan Zhang, Kaichen Zhang, Bo Li, Guangtao Zeng, Jingkang Yang, Yuanhan Zhang, Ziyue Wang, Haoran Tan, Chunyuan Li, and Ziwei Liu · 2024
Later among the works it cites.
Llava-next: A strong zero-shot video understanding model, April 2024
Yuanhan Zhang, Bo Li, haotian Liu, Yong jae Lee, Liangke Gui, Di Fu, Jiashi Feng, Ziwei Liu, and Chunyuan Li · 2024
Later among the works it cites.
Mlvu: A comprehensive benchmark for multi-task long video understanding
Junjie Zhou, Yan Shu, Bo Zhao, Boya Wu, Shitao Xiao, Xi Yang, Yongping Xiong, Bo Zhang, Tiejun Huang, and Zheng Liu · 2024
Later among the works it cites.
Membership inference attacks against vision-language models
Yuke Hu, Zheng Li, Zhihao Liu, Yang Zhang, Zhan Qin, Kui Ren, and Chun Chen · 2025
Closest in time.
Videoeval-pro: Robust and realistic long video understanding evaluation
Wentao Ma, Weiming Ren, Yiming Jia, Zhuofeng Li, Ping Nie, Ge Zhang, and Wenhu Chen · 2025
Closest in time.
Visual text processing: A comprehensive review and unified evaluation
Yan Shu, Weichao Zeng, Fangmin Zhao, Zeyu Chen, Zhenhang Li, Xiaomeng Yang, Yu Zhou, Paolo Rota, Xiang Bai, Lianwen Jin, et al · 2025
Closest in time.
Membership inference attacks on large-scale models: A survey
Hengyu Wu and Yang Cao · 2025
Closest in time.
Memory-enhanced retrieval augmentation for long video understanding
Huaying Yuan, Zheng Liu, Minhao Qin, Hongjin Qian, Y Shu, Zhicheng Dou, and Ji-Rong Wen · 2025
Closest in time.