Fetching the paper…
Reading the bibliography…
Multimodal LLMs have advanced vision-language tasks but still struggle with understanding video scenes.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang · 2017
Earlier work this paper cites.
Action genome: Actions as compositions of spatio-temporal scene graphs
Jingwei Ji, Ranjay Krishna, Li Fei-Fei, and Juan Carlos Niebles · 2020
Earlier work this paper cites.
Hypergraph attention networks for multimodal learning
Eun-Sol Kim, Woo Young Kang, Kyoung-Woon On, Yu-Jung Heo, and Byoung-Tak Zhang · 2020
Earlier work this paper cites.
Hypergraph convolution and hypergraph attention
Song Bai, Feihu Zhang, and Philip HS Torr · 2021
Earlier work this paper cites.
Spatial-temporal transformer for dynamic scene graph generation
Yuren Cong, Wentong Liao, Hanno Ackermann, Bodo Rosenhahn, and Michael Ying Yang · 2021
Earlier work this paper cites.
Spatial-temporal transformer for dynamic scene graph generation
Yuren Cong, Wentong Liao, H. Ackermann, M. Yang, and B. Rosenhahn · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Earlier work this paper cites.
Exploiting long-term dependencies for generating dynamic scene graphs
Shengyu Feng, Hesham Mostafa, Marcel Nassar, Somdeb Majumdar, and Subarna Tripathi · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Video visual relation detection via iterative inference
Xindi Shang, Yicong Li, Junbin Xiao, Wei Ji, and Tat-Seng Chua · 2021
Earlier work this paper cites.
Target adaptive context aggregation for video scene graph generation
Yao Teng, Limin Wang, Zhifeng Li, and Gangshan Wu · 2021
Earlier work this paper cites.
Hierarchical memory learning for fine-grained scene graph generation
Youming Deng, Yansheng Li, Yongjun Zhang, Xiang Xiang, Jian Wang, Jingdong Chen, and Jiayi Ma · 2022
Earlier work this paper cites.
Hgnn+: General hypergraph neural networks
Yue Gao, Yifan Feng, Shuyi Ji, and Rongrong Ji · 2022
Earlier work this paper cites.
Towards open-vocabulary scene graph generation with prompt-based finetuning
Tao He, Lianli Gao, Jingkuan Song, and Yuan-Fang Li · 2022
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Earlier work this paper cites.
Sgtr: End-to-end scene graph generation with transformer
Rongjie Li, Songyang Zhang, and Xuming He · 2022
Earlier work this paper cites.
Multi-hyperedge hypergraph for group activity recognition
Wanxin Li, Wei Xie, Zhigang Tu, Wei Wang, and Lianghao Jin · 2022
Earlier work this paper cites.
End-to-end generative pretraining for multimodal video captioning
Paul Hongsuck Seo, Arsha Nagrani, Anurag Arnab, and Cordelia Schmid · 2022
Earlier work this paper cites.
Modeling semantic composition with syntactic hypergraph for video question answering
Zenan Xu, Wanjun Zhong, Qinliang Su, Zijing Ou, and Fuwei Zhang · 2022
Cited alongside, same era.
Panoptic scene graph generation
Jingkang Yang, Yi Zhe Ang, Zujin Guo, Kaiyang Zhou, Wayne Zhang, and Ziwei Liu · 2022
Cited alongside, same era.
Deep hypergraph structure learning
Zizhao Zhang, Yifan Feng, Shihui Ying, and Yue Gao · 2022
Cited alongside, same era.
More knowledge, less bias: Unbiasing scene graph generation with explicit ontological adjustment
Zhanwen Chen, Saed Rezayi, and Sheng Li · 2023
Cited alongside, same era.
Reltr: Relation transformer for scene graph generation
Yuren Cong, Michael Ying Yang, and Bodo Rosenhahn · 2023
Cited alongside, same era.
Dsgg: Dense relation transformer for an end-to-end scene graph generation
Zeeshan Hayder and Xuming He · 2024
Closest in time.
Ma-lmm: Memory-augmented large multimodal model for long-term video understanding
Bo He, Hengduo Li, Young Kyun Jang, Menglin Jia, Xuefei Cao, Ashish Shah, Abhinav Shrivastava, and Ser-Nam Lim · 2024
Closest in time.
Egtr: Extracting graph from transformer for scene graph generation
Jinbae Im, JeongYeon Nam, Nokyung Park, Hyungmin Lee, and Seunghyun Park · 2024
Closest in time.
Chat-univi: Unified visual representation empowers large language models with image and video understanding
Peng Jin, Ryuichi Takanobu, Wancai Zhang, Xiaochun Cao, and Li Yuan · 2024
Closest in time.
Llm4sgg: Large language models for weakly supervised scene graph generation
Kibum Kim, Kanghoon Yoon, Jaehyeong Jeon, Yeonjun In, Jinyoung Moon, Donghyun Kim, and Chanyoung Park · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scenegenie: Scene graph guided diffusion models for image synthesis
Azade Farshad, Yousef Yeganeh, Yu Chi, Chengzhi Shen, Böjrn Ommer, and Nassir Navab · 2023
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Cited alongside, same era.
Fast contextual scene graph generation with unbiased context augmentation
Tianlei Jin, Fangtai Guo, Qiwei Meng, Shiqiang Zhu, Xiangming Xi, Wen Wang, Zonghao Mu, and Wei Song · 2023
Cited alongside, same era.
Meltr: Meta loss transformer for learning to fine-tune video foundation models
Dohwan Ko, Joonmyung Choi, Hyeong Kyu Choi, Kyoung-Woon On, Byungseok Roh, and Hyunwoo J Kim · 2023
Cited alongside, same era.
Is-ggt: Iterative scene graph generation with generative transformers
Sanjoy Kundu and Sathyanarayanan N Aakur · 2023
Cited alongside, same era.
Video-llava: Learning united visual representation by alignment before projection
Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, and Li Yuan · 2023
Cited alongside, same era.
Unbiased scene graph generation in videos
Sayak Nag, Kyle Min, Subarna Tripathi, and Amit K Roy-Chowdhury · 2023
Cited alongside, same era.
From pixels to graphs: Open-vocabulary scene graph generation with vision-language models
Rongjie Li, Songyang Zhang, Dahua Lin, Kai Chen, and Xuming He · 2024
Closest in time.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2024
Closest in time.
Video-ChatGPT: Towards detailed video understanding via large vision and language models
Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Khan · 2024
Closest in time.
Hig: Hierarchical interlacement graph approach to scene graph generation in video understanding
Trong-Thuan Nguyen, Pha Nguyen, and Khoa Luu · 2024
Closest in time.
CYCLO: Cyclic graph transformer approach to multi-object relationship modeling in aerial videos
Trong-Thuan Nguyen, Pha Nguyen, Li Xin, Cothren Jackson, Yilmaz Alper, and Khoa Luu · 2024
Closest in time.
Towards scene graph anticipation
Rohith Peddi, Saksham Singh, Parag Singla, Vibhav Gogate, et al · 2024
Closest in time.
Moviechat: From dense token to sparse memory for long video understanding
Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo, Tian Ye, Yanting Zhang, et al · 2024
Closest in time.
Graph (graph): A nested graph-based framework for early accident anticipation
Nupur Thakur, PrasanthSai Gouripeddi, and Baoxin Li · 2024
Closest in time.
Multi-object event graph representation learning for video question answering
Yanan Wang, Shuichiro Haruta, Donghuo Zeng, Julio Vizcarra, and Mori Kurokawa · 2024
Closest in time.
Sportshhi: A dataset for human-human interaction detection in sports videos
Tao Wu, Runyu He, Gangshan Wu, and Limin Wang · 2024
Closest in time.
Commonscenes: Generating commonsense 3d indoor scenes with scene graphs
Guangyao Zhai, Evin Pınar Örnek, Shun-Cheng Wu, Yan Di, Federico Tombari, Nassir Navab, and Benjamin Busam · 2024
Closest in time.
Dynamical attention hypergraph convolutional network for group activity recognition
Xiaolin Zhu, Dongli Wang, Jianxun Li, Rui Su, Qin Wan, and Yan Zhou · 2024
Closest in time.
Llama-vid: An image is worth 2 tokens in large language models
Yanwei Li, Chengyao Wang, and Jiaya Jia · 2025
Closest in time.