Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have allowed recent LLM-based approaches to achieve excellent performance on long-video understanding benchmarks.
The representation and recognition of action using temporal templates
James Davis and Aaron Bobick · 1997
Earlier work this paper cites.
Pattern recognition and machine learning (information science and statistics), 2006
Christopher M. Bishop · 2006
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2015
Earlier work this paper cites.
Dynamic image networks for action recognition
Hakan Bilen, Basura Fernando, Efstratios Gavves, Andrea Vedaldi, and Stephen Gould · 2016
Earlier work this paper cites.
The “something something” video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fründ, Peter N. Yianilos, Moritz Mueller-Freitag, Florian Hoppe, Christian Thurau, Ingo Bax, and Roland Memisevic · 2017
Earlier work this paper cites.
Alternative semantic representations for zero-shot human action recognition
Qian Wang and Ke Chen · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang · 2017
Earlier work this paper cites.
Leveraging video descriptions to learn video question answering
Kuo-Hao Zeng, Tseng-Hung Chen, Ching-Yao Chuang, Yuan-Hong Liao, Juan Carlos Niebles, and Min Sun · 2017
Earlier work this paper cites.
Single image action recognition using semantic body part actions
Zhichen Zhao, Huimin Ma, and Shaodi You · 2017
Earlier work this paper cites.
TVQA: Localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara Berg · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training, 2018
Alec Radford and Karthik Narasimhan · 2018
Earlier work this paper cites.
Combined static and motion features for deep-networks-based activity recognition in videos
Sameera Ramasinghe, Jathushan Rajasegaran, Vinoj Jayasundara, Kanchana Ranasinghe, Ranga Rodrigo, and Ajith A. Pasqual · 2018
Earlier work this paper cites.
Fast online object tracking and segmentation: A unifying approach
Qiang Wang, Li Zhang, Luca Bertinetto, Weiming Hu, and Philip H. S. Torr · 2018
Earlier work this paper cites.
Still image action recognition by predicting spatial-temporal pixel evolution
Marjaneh Safaei and Hassan Foroosh · 2019
Earlier work this paper cites.
Activitynet-qa: A dataset for understanding complex web videos via question answering
Zhou Yu, D. Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao · 2019
Earlier work this paper cites.
TVQA+: Spatio-temporal grounding for video question answering
Jie Lei, Licheng Yu, Tamara Berg, and Mohit Bansal · 2020
Earlier work this paper cites.
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Self-supervised video transformer
Kanchana Ranasinghe, Muzammal Naseer, Salman Hameed Khan, Fahad Shahbaz Khan, and Michael S. Ryoo · 2021
Earlier work this paper cites.
NExT-QA: Next phase of question-answering to explaining temporal actions
Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua · 2021
Earlier work this paper cites.
Just ask: Learning to answer questions from millions of narrated videos
Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid · 2021
Earlier work this paper cites.
Revisiting the “video” in video-language understanding
S. Buch, Cristobal Eyzaguirre, Adrien Gaidon, Jiajun Wu, Li Fei-Fei, and Juan Carlos Niebles · 2022
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Earlier work this paper cites.
Faithful reasoning using large language models
Antonia Creswell and Murray Shanahan · 2022
Earlier work this paper cites.
An empirical study of end-to-end video-language transformers with masked visual modeling
Tsu-Jui Fu, Linjie Li, Zhe Gan, Kevin Lin, William Yang Wang, Lijuan Wang, and Zicheng Liu · 2022
Cited alongside, same era.
Vtc: Improving video-text retrieval with user comments
Laura Hanu, James Thewlis, Yuki M. Asano, and C. Rupprecht · 2022
Cited alongside, same era.
Knowledge-augmented language models for cause-effect relation classification
Pedram Hosseini, David A. Broniatowski, and Mona Diab · 2022
Cited alongside, same era.
Simple open-vocabulary object detection with vision transformers
Matthias Minderer, Alexey A. Gritsenko, Austin Stone, Maxim Neumann, Dirk Weissenborn, Alexey Dosovitskiy, Aravindh Mahendran, Anurag Arnab, Mostafa Dehghani, Zhuoran Shen, Xiao Wang, Xiaohua Zhai, Thomas Kipf, and Neil Houlsby · 2022
Cited alongside, same era.
Internvideo: General video foundation models via generative and discriminative learning
Leveraging large language models for multiple choice question answering
Joshua Robinson, Christopher Rytting, and David Wingate · 2023
Later among the works it cites.
Language models are causal knowledge extractors for zero-shot video question answering
Hung-Ting Su, Yulei Niu, Xudong Lin, Winston H. Hsu, and Shih-Fu Chang · 2023
Later among the works it cites.
Vipergpt: Visual inference via python execution for reasoning
Dídac Surís, Sachit Menon, and Carl Vondrick · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang, Jilan Xu, Yi Liu, Zun Wang, et al · 2022
Cited alongside, same era.
Videococa: Video-text modeling with zero-shot transfer from contrastive captioners
Shen Yan, Tao Zhu, Zirui Wang, Yuan Cao, Mi Zhang, Soham Ghosh, Yonghui Wu, and Jiahui Yu · 2022
Cited alongside, same era.
Zero-shot video question answering via frozen bidirectional language models
Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid · 2022
Cited alongside, same era.
Time is MattEr: Temporal self-supervision for video transformers
Sukmin Yun, Jaehyung Kim, Dongyoon Han, Hwanjun Song, Jung-Woo Ha, and Jinwoo Shin · 2022
Cited alongside, same era.
Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language
Andy Zeng, Adrian Wong, Stefan Welker, Krzysztof Choromanski, Federico Tombari, Aveek Purohit, Michael Ryoo, Vikas Sindhwani, Johnny Lee, Vincent Vanhoucke, et al · 2022
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2023
Cited alongside, same era.
Language as the medium: Multimodal video classification through text only
Laura Hanu, Anita Lilla Vero, and James Thewlis · 2023
Cited alongside, same era.
Albert Qiaochu Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L’elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Cited alongside, same era.
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Gemini in reasoning: Unveiling commonsense in multimodal large language models
Yuqing Wang and Yun Zhao · 2023
Later among the works it cites.
System 2 attention (is something you might need too)
Jason Weston and Sainbayar Sukhbaatar · 2023
Later among the works it cites.
Contrastive video question answering via video graph transformer
Junbin Xiao, Pan Zhou, Angela Yao, Yicong Li, Richang Hong, Shuicheng Yan, and Tat-Seng Chua · 2023
Later among the works it cites.
Large language models as commonsense knowledge for large-scale task planning, 2023
Zirui Zhao, Wee Sun Lee, and David Hsu · 2023
Later among the works it cites.
Phi-3 technical report: A highly capable language model locally on your phone
Marah Abdin et al · 2024
Closest in time.
Investigating cultural alignment of large language models, 2024
Badr AlKhamissi, Muhammad ElNokrashy, Mai AlKhamissi, and Mona Diab · 2024
Closest in time.
Memory consolidation enables long-context video understanding
Ivana Balavzevi’c, Yuge Shi, Pinelopi Papalampidi, Rahma Chaabouni, Skanda Koppula, and Olivier J. H’enaff · 2024
Closest in time.
Language repository for long video understanding
Kumara Kahatapitiya, Kanchana Ranasinghe, Jongwoo Park, and Michael S Ryoo · 2024
Closest in time.
Mvbench: A comprehensive multi-modal video understanding benchmark
Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, Limin Wang, and Yu Qiao · 2024
Closest in time.
Llara: Supercharging robot learning data for vision-language policy
Xiang Li, Cristina Mata, Jongwoo Park, Kumara Kahatapitiya, Yoo Sung Jang, Jinghuan Shang, Kanchana Ranasinghe, Ryan Burgert, Mu Cai, Yong Jae Lee, and Michael S. Ryoo · 2024
Closest in time.
Morevqa: Exploring modular reasoning models for video question answering
Juhong Min, Shyamal Buch, Arsha Nagrani, Minsu Cho, and Cordelia Schmid · 2024
Closest in time.
Too many frames, not all useful: Efficient strategies for long-form video qa
Jong Sung Park, Kanchana Ranasinghe, Kumara Kahatapitiya, Wonjeong Ryoo, Donghyun Kim, and Michael S. Ryoo · 2024
Closest in time.
Learning to localize objects improves spatial reasoning in visual-llms
Kanchana Ranasinghe, Satya Narayan Shukla, Omid Poursaeed, Michael S Ryoo, and Tsung-Yu Lin · 2024
Closest in time.
Longvideobench: A benchmark for long-context interleaved video-language understanding, 2024
Haoning Wu, Dongxu Li, Bei Chen, and Junnan Li · 2024
Closest in time.
Can i trust your answer? visually grounded video question answering
Junbin Xiao, Angela Yao, Yicong Li, and Tat-Seng Chua · 2024
Closest in time.
Penetrative ai: Making llms comprehend the physical world, 2024
Huatao Xu, Liying Han, Qirui Yang, Mo Li, and Mani Srivastava · 2024
Closest in time.
Self-chained image-language model for video localization and question answering
Shoubin Yu, Jaemin Cho, Prateek Yadav, and Mohit Bansal · 2024
Closest in time.
Videoagent: Long-form video understanding with large language model as agent
Xiaohan Wang, Yuhui Zhang, Orr Zohar, and Serena Yeung-Levy · 2025
Closest in time.