Fetching the paper…
Reading the bibliography…
In this paper, we introduce Motion-Grounded Video Reasoning, a new motion understanding task that requires generating visual answers (video segmentation masks) according to the input question, and hence needs implicit spatiotemporal reasoning and grounding.
The distribution of the flora in the alpine zone. 1
Paul Jaccard · 1912
Earlier work this paper cites.
Measures of the amount of ecologic association between species
Lee R Dice · 1945
Earlier work this paper cites.
Human motion analysis: A review
Jake K Aggarwal and Quin Cai · 1999
Earlier work this paper cites.
Understanding purposeful human motion
Christopher Richard Wren and Alex P Pentland · 1999
Earlier work this paper cites.
Implicit human computer interaction through context
Albrecht Schmidt · 2000
Earlier work this paper cites.
Grounding language in action
Arthur M Glenberg and Michael P Kaschak · 2002
Earlier work this paper cites.
Human motion: Modeling and recognition of actions and interactions
Jake K Aggarwal and Sangho Park · 2004
Earlier work this paper cites.
Combining appearance and structure from motion features for road scene understanding
Paul Sturgess, Karteek Alahari, Lubor Ladicky, and Philip HS Torr · 2009
Earlier work this paper cites.
Scene understanding by statistical modeling of motion patterns
Imran Saleemi, Lance Hartung, and Mubarak Shah · 2010
Earlier work this paper cites.
Real-time indoor scene understanding using bayesian filtering with motion cues
Grace Tsai, Changhai Xu, Jingen Liu, and Benjamin Kuipers · 2011
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Hierarchical aligned cluster analysis for temporal clustering of human motion
Feng Zhou, Fernando De la Torre, and Jessica K Hodgins · 2012
Earlier work this paper cites.
Grounding action descriptions in videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal · 2013
Earlier work this paper cites.
THUMOS challenge: Action recognition with a large number of classes
Y.-G. Jiang, J. Liu, A. Roshan Zamir, G. Toderici, I. Laptev, M. Shah, and R. Sukthankar · 2014
Earlier work this paper cites.
ReferItGame: Referring to objects in photographs of natural scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Deepdriving: Learning affordance for direct perception in autonomous driving
Chenyi Chen, Ari Seff, Alain Kornhauser, and Jianxiong Xiao · 2015
Earlier work this paper cites.
Finding action tubes
Georgia Gkioxari and Jitendra Malik · 2015
Earlier work this paper cites.
Newtonian scene understanding: Unfolding the dynamics of objects in static images
Roozbeh Mottaghi, Hessam Bagherinezhad, Mohammad Rastegari, and Ali Farhadi · 2016
Earlier work this paper cites.
Movieqa: Understanding stories in movies through question-answering
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
The” something something” video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Earlier work this paper cites.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim · 2017
Earlier work this paper cites.
Object level visual reasoning in videos
Fabien Baradel, Natalia Neverova, Christian Wolf, Julien Mille, and Greg Mori · 2018
Earlier work this paper cites.
End-to-end learning of motion representation for video understanding
Lijie Fan, Wenbing Huang, Chuang Gan, Stefano Ermon, Boqing Gong, and Junzhou Huang · 2018
Earlier work this paper cites.
Ava: A video dataset of spatio-temporally localized atomic visual actions
Chunhui Gu, Chen Sun, David A Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, et al · 2018
Cited alongside, same era.
What makes a video a video: Analyzing temporal information in video understanding models and datasets
De-An Huang, Vignesh Ramanathan, Dhruv Mahajan, Lorenzo Torresani, Manohar Paluri, Li Fei-Fei, and Juan Carlos Niebles · 2018
Cited alongside, same era.
Real-world anomaly detection in surveillance videos
Waqas Sultani, Chen Chen, and Mubarak Shah · 2018
Cited alongside, same era.
Tube-cnn: Modeling temporal evolution of appearance for object detection in video
Tuan-Hung Vu, Anton Osokin, and Ivan Laptev · 2018
Cited alongside, same era.
Youtube-vos: Sequence-to-sequence video object segmentation
Ning Xu, Linjie Yang, Yuchen Fan, Jianchao Yang, Dingcheng Yue, Yuchen Liang, Brian Price, Scott Cohen, and Thomas Huang · 2018
Cited alongside, same era.
Lisa: Reasoning segmentation via large language model
Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia · 2023
Later among the works it cites.
Robust referring video object segmentation with cyclic structural consensus
Xiang Li, Jinglu Wang, Xiaohao Xu, Xiao Li, Bhiksha Raj, and Yan Lu · 2023
Later among the works it cites.
Collaborative static and dynamic vision-language streams for spatio-temporal video grounding
Zihang Lin, Chaolei Tan, Jian-Fang Hu, Zhi Jin, Tiancai Ye, and Wei-Shi Zheng · 2023
Later among the works it cites.
Video-chatgpt: Towards detailed video understanding via large vision and language models
Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan · 2023
Later among the works it cites.
Spectrum-guided multi-granularity referring video object segmentation
Bo Miao, Mohammed Bennamoun, Yongsheng Gao, and Ajmal Mian · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Video object segmentation with language referring expressions
Anna Khoreva, Anna Rohrbach, and Bernt Schiele · 2019
Cited alongside, same era.
A review of tracking, prediction and decision making methods for autonomous driving
Florin Leon and Marius Gavrilescu · 2019
Cited alongside, same era.
Activitynet-qa: A dataset for understanding complex web videos via question answering
Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao · 2019
Cited alongside, same era.
Context-aware human motion prediction
Enric Corona, Albert Pumarola, Guillem Alenya, and Francesc Moreno-Noguer · 2020
Cited alongside, same era.
Learning where to focus for efficient video object detection
Zhengkai Jiang, Yu Liu, Ceyuan Yang, Jihao Liu, Peng Gao, Qian Zhang, Shiming Xiang, and Chunhong Pan · 2020
Cited alongside, same era.
Urvos: Unified referring video object segmentation network with a large-scale benchmark
Seonguk Seo, Joon-Young Lee, and Bohyung Han · 2020
Cited alongside, same era.
Finegym: A hierarchical video dataset for fine-grained action understanding
Dian Shao, Yue Zhao, Bo Dai, and Dahua Lin · 2020
Cited alongside, same era.
Later among the works it cites.
Pg-video-llava: Pixel grounding large video-language models
Shehan Munasinghe, Rusiru Thushara, Muhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Mubarak Shah, and Fahad Khan · 2023
Later among the works it cites.
Segment every reference object in spatial and temporal spaces
Jiannan Wu, Yi Jiang, Bin Yan, Huchuan Lu, Zehuan Yuan, and Ping Luo · 2023
Later among the works it cites.
A simple llm framework for long-range video question-answering
Ce Zhang, Taixi Lu, Md Mohaiminul Islam, Ziyang Wang, Shoubin Yu, Mohit Bansal, and Gedas Bertasius · 2023
Later among the works it cites.
Pixellm: Pixel reasoning with large multimodal model
Ren Zhongwei, Huang Zhicheng, Wei Yunchao, Zhao Yao, Fu Dongmei, Feng Jiashi, and Jin Xiaojie · 2023
Later among the works it cites.
One token to seg them all: Language instructed reasoning segmentation in videos
Zechen Bai, Tong He, Haiyang Mei, Pichao Wang, Ziteng Gao, Joya Chen, Lei Liu, Zheng Zhang, and Mike Zheng Shou · 2024
Closest in time.
Context-guided spatio-temporal video grounding
Xin Gu, Heng Fan, Yan Huang, Tiejian Luo, and Libo Zhang · 2024
Closest in time.
Decoupling static and hierarchical motion perception for referring video segmentation
Shuting He and Henghui Ding · 2024
Closest in time.
Lita: Language instructed temporal-localization assistant
De-An Huang, Shijia Liao, Subhashree Radhakrishnan, Hongxu Yin, Pavlo Molchanov, Zhiding Yu, and Jan Kautz · 2024
Closest in time.
Mvbench: A comprehensive multi-modal video understanding benchmark
Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al · 2024
Closest in time.
Vila: On pre-training for visual language models
Ji Lin, Hongxu Yin, Wei Ping, Pavlo Molchanov, Mohammad Shoeybi, and Song Han · 2024
Closest in time.
Towards temporally consistent referring video object segmentation
Bo Miao, Mohammed Bennamoun, Yongsheng Gao, Mubarak Shah, and Ajmal Mian · 2024
Closest in time.
Perception test: A diagnostic benchmark for multimodal video models
Viorica Patraucean, Lucas Smaira, Ankush Gupta, Adria Recasens, Larisa Markeeva, Dylan Banarse, Skanda Koppula, Mateusz Malinowski, Yi Yang, Carl Doersch, et al · 2024
Closest in time.
Glamm: Pixel grounding large multimodal model
Hanoona Rasheed, Muhammad Maaz, Sahal Shaji, Abdelrahman Shaker, Salman Khan, Hisham Cholakkal, Rao M. Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S. Khan · 2024
Closest in time.
Sam 2: Segment anything in images and videos
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Dollár, and Christoph Feichtenhofer · 2024
Closest in time.
Can i trust your answer? visually grounded video question answering
Junbin Xiao, Angela Yao, Yicong Li, and Tat-Seng Chua · 2024
Closest in time.
Visa: Reasoning video object segmentation via large language models
Cilin Yan, Haochen Wang, Shilin Yan, Xiaolong Jiang, Yao Hu, Guoliang Kang, Weidi Xie, and Efstratios Gavves · 2024
Closest in time.
Crema: Multimodal compositional video reasoning via efficient modular adaptation and fusion
Shoubin Yu, Jaehong Yoon, and Mohit Bansal · 2024
Closest in time.
Villa: Video reasoning segmentation with large language model
Rongkun Zheng, Lu Qi, Xi Chen, Yi Wang, Kun Wang, Yu Qiao, and Hengshuang Zhao · 2024
Closest in time.
Llama-vid: An image is worth 2 tokens in large language models
Yanwei Li, Chengyao Wang, and Jiaya Jia · 2025
Closest in time.