Fetching the paper…
Reading the bibliography…
Video Large Language Models (Vid-LLMs) have made remarkable advancements in comprehending video content for QA dialogue.
Adam: A method for stochastic optimization
Diederik P Kingma · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Tall: Temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Earlier work this paper cites.
Temporal action detection with structured segment networks
Yue Zhao, Yuanjun Xiong, Limin Wang, Zhirong Wu, Xiaoou Tang, and Dahua Lin · 2017
Earlier work this paper cites.
Temporal segment networks for action recognition in videos
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2018
Earlier work this paper cites.
Weakly supervised video moment retrieval from text queries
Niluthpol Chowdhury Mithun, Sujoy Paul, and Amit K Roy-Chowdhury · 2019
Earlier work this paper cites.
Progressive bilateral-context driven model for post-processing person re-identification
Min Cao, Chen Chen, Hao Dou, Xiyuan Hu, Silong Peng, and Arjan Kuijper · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy · 2020
Earlier work this paper cites.
Tea: Temporal excitation and aggregation for action recognition
Yan Li, Bin Ji, Xintian Shi, Jianguo Zhang, Bin Kang, and Limin Wang · 2020
Earlier work this paper cites.
End-to-end multi-modal video temporal grounding
Yi-Wen Chen, Yi-Hsuan Tsai, and Ming-Hsuan Yang · 2021
Earlier work this paper cites.
Fast video moment retrieval
Junyu Gao and Changsheng Xu · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al · 2021
Earlier work this paper cites.
Detecting moments and highlights in videos via natural language queries
Jie Lei, Tamara L Berg, and Mohit Bansal · 2021
Earlier work this paper cites.
Interventional video grounding with dual contrastive learning
Guoshun Nan, Rui Qiao, Yao Xiao, Jun Liu, Sicong Leng, Hao Zhang, and Wei Lu · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Sutd-trafficqa: A question answering benchmark and an efficient network for video reasoning over traffic events
Li Xu, He Huang, and Jun Liu · 2021
Earlier work this paper cites.
Visual place recognition via local affine preserving matching
Xinyu Ye and Jiayi Ma · 2021
Earlier work this paper cites.
Image-text retrieval: A survey on recent research and development
Min Cao, Shiping Li, Juntao Li, Liqiang Nie, and Min Zhang · 2022
Earlier work this paper cites.
Learning semantic-aligned feature representation for text-based person search
Shiping Li, Min Cao, and Min Zhang · 2022
Cited alongside, same era.
Meta spatio-temporal debiasing for video scene graph generation
Li Xu, Haoxuan Qu, Jason Kuen, Jiuxiang Gu, and Jun Liu · 2022
Cited alongside, same era.
Neighborhood manifold preserving matching for visual place recognition
Xinyu Ye and Jiayi Ma · 2022
Cited alongside, same era.
Instructdet: Diversifying referring object detection with generalized instructions
Ronghao Dang, Jiangyan Feng, Haodong Zhang, Chongjian Ge, Lin Song, Lijun Gong, Chengju Liu, Qijun Chen, Feng Zhu, Rui Zhao, et al · 2023
Cited alongside, same era.
System-status-aware adaptive network for online streaming video understanding
Lin Geng Foo, Jia Gong, Zhipeng Fan, and Jun Liu · 2023
Cited alongside, same era.
Multi-scale spatial-temporal attention networks for functional connectome classification
Youyong Kong, Xiaotong Zhang, Wenhan Wang, Yue Zhou, Yueying Li, and Yonggui Yuan · 2024
Closest in time.
Collavo: Crayon large language and vision model
Byung-Kwan Lee, Beomchan Park, Chae Won Kim, and Yong Man Ro · 2024
Closest in time.
Hello gpt-4o
OpenAI · 2024
Closest in time.
Timechat: A time-sensitive multimodal large language model for long video understanding
Shuhuai Ren, Linli Yao, Shicheng Li, Xu Sun, and Lu Hou · 2024
Closest in time.
Towards more unified in-context visual understanding
Dianmo Sheng, Dongdong Chen, Zhentao Tan, Qiankun Liu, Qi Chu, Jianmin Bao, Tao Gong, Bin Liu, Shengwei Xu, and Nenghai Yu · 2024
Closest in time.
Add-it: Training-free object insertion in images with pretrained diffusion models, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Univtg: Towards unified video-language temporal grounding
Kevin Qinghong Lin, Pengchuan Zhang, Joya Chen, Shraman Pramanick, Difei Gao, Alex Jinpeng Wang, Rui Yan, and Mike Zheng Shou · 2023
Cited alongside, same era.
Zero-shot model diagnosis
Jinqi Luo, Zhaoning Wang, Chen Henry Wu, Dong Huang, and Fernando De la Torre · 2023
Cited alongside, same era.
Video-chatgpt: Towards detailed video understanding via large vision and language models
Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan · 2023
Cited alongside, same era.
What does clip know about a red circle? visual prompt engineering for vlms
Aleksandar Shtedritski, Christian Rupprecht, and Andrea Vedaldi · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Cited alongside, same era.
Visual-semantic network: a visual and semantic enhanced model for gesture recognition
Yizhe Wang, Congqi Cao, and Yanning Zhang · 2023
Cited alongside, same era.
Vqne: Variational quantum network embedding with application to network alignment
Xinyu Ye, Ge Yan, and Junchi Yan · 2023
Cited alongside, same era.
Yoad Tewel, Rinon Gal, Dvir Samuel, Yuval Atzmon, Lior Wolf, and Gal Chechik · 2024
Closest in time.
Towards open-world grasping with large vision-language models
Georgios Tziafas and Hamidreza Kasaei · 2024
Closest in time.
Omniedit: Building image editing generalist models through specialist supervision, 2024
Cong Wei, Zheyang Xiong, Weiming Ren, Xinrun Du, Ge Zhang, and Wenhu Chen · 2024
Closest in time.
A glance at in-context learning
Yongliang Wu and Xu Yang · 2024
Closest in time.
Exploring diverse in-context configurations for image captioning
Xu Yang, Yongliang Wu, Mingzhuo Yang, Haokun Chen, and Xin Geng · 2024
Closest in time.
Cpt: Colorful prompt tuning for pre-trained vision-language models
Yuan Yao, Ao Zhang, Zhengyan Zhang, Zhiyuan Liu, Tat-Seng Chua, and Maosong Sun · 2024
Closest in time.
Bridge the modality and capability gaps in vision-language model selection
Chao Yi, Yuhang He, De-Chuan Zhan, and Han-Jia Ye · 2024
Closest in time.
Dataset regeneration for sequential recommendation
Mingjia Yin, Hao Wang, Wei Guo, Yong Liu, Suojuan Zhang, Sirui Zhao, Defu Lian, and Enhong Chen · 2024
Closest in time.
Multi-modal in-context learning makes an ego-evolving scene text recognizer
Zhen Zhao, Jingqun Tang, Chunhui Lin, Binghong Wu, Can Huang, Hao Liu, Xin Tan, Zhizhong Zhang, and Yuan Xie · 2024
Closest in time.
Minedreamer: Learning to follow instructions via chain-of-imagination for simulated-world control
Enshen Zhou, Yiran Qin, Zhenfei Yin, Yuzhou Huang, Ruimao Zhang, Lu Sheng, Yu Qiao, and Jing Shao · 2024
Closest in time.
Homeomorphism prior for false positive and negative problem in medical image dense contrastive representation learning
Yuting He, Boyu Wang, Rongjun Ge, Yang Chen, Guanyu Yang, and Shuo Li · 2025
Closest in time.
Llava-st: A multimodal large language model for fine-grained spatial-temporal understanding
Hongyu Li, Jinyu Chen, Ziyu Wei, Shaofei Huang, Tianrui Hui, Jialin Gao, Xiaoming Wei, and Si Liu · 2025
Closest in time.
Groma: Localized visual tokenization for grounding multimodal large language models
Chuofan Ma, Yi Jiang, Jiannan Wu, Zehuan Yuan, and Xiaojuan Qi · 2025
Closest in time.
Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl
Yingzhe Peng, Gongrui Zhang, Miaosen Zhang, Zhiyuan You, Jie Liu, Qipeng Zhu, Kai Yang, Xingzhong Xu, Xin Geng, and Xu Yang · 2025
Closest in time.