Fetching the paper…
Reading the bibliography…
Recent breakthroughs in Multimodal Large Language Models (MLLMs) have gained significant recognition within the deep learning community, where the fusion of the Video Foundation Models (VFMs) and Large Language Models(LLMs) has proven instrumental in constructing robust video understanding systems, effectively surmounting constraints associated with predefined visual tasks.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
The thumos challenge on action recognition for videos “in the wild”
Haroon Idrees, Amir R Zamir, Yu-Gang Jiang, Alex Gorban, Ivan Laptev, Rahul Sukthankar, and Mubarak Shah · 2017
Earlier work this paper cites.
Untrimmednets for weakly supervised action recognition and detection
Limin Wang, Yuanjun Xiong, Dahua Lin, and Luc Van Gool · 2017
Earlier work this paper cites.
Temporal action detection with structured segment networks
Yue Zhao, Yuanjun Xiong, Limin Wang, Zhirong Wu, Xiaoou Tang, and Dahua Lin · 2017
Earlier work this paper cites.
Rethinking the faster r-cnn architecture for temporal action localization
Yu Wei Chao, Sudheendra Vijayanarasimhan, Bryan Seybold, David A. Ross, Jia Deng, and Rahul Sukthankar · 2018
Earlier work this paper cites.
Weakly supervised action localization by sparse temporal pooling network
Phuc Nguyen, Ting Liu, Gautam Prasad, and Bohyung Han · 2018
Earlier work this paper cites.
W-talc: Weakly-supervised temporal activity localization and classification
Sujoy Paul, Sourya Roy, and Amit K Roy-Chowdhury · 2018
Earlier work this paper cites.
Graph convolutional networks for temporal action localization
Runhao Zeng, Wenbing Huang, Mingkui Tan, Yu Rong, Peilin Zhao, Junzhou Huang, and Chuang Gan · 2019
Earlier work this paper cites.
Boundary content graph neural network for temporal action proposal generation
Yueran Bai, Yingying Wang, Yunhai Tong, Yang Yang, Qiyue Liu, and Junhui Liu · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
G-tad: Sub-graph localization for temporal action detection
Mengmeng Xu, Chen Zhao, David S Rojas, Ali Thabet, and Bernard Ghanem · 2020
Earlier work this paper cites.
Two-stream consensus network for weakly-supervised temporal action localization
Yuanhao Zhai, Le Wang, Wei Tang, Qilin Zhang, Junsong Yuan, and Gang Hua · 2020
Earlier work this paper cites.
Cross-modal consensus network for weakly supervised temporal action localization
Fa-Ting Hong, Jia-Chang Feng, Dan Xu, Ying Shan, and Wei-Shi Zheng · 2021
Cited alongside, same era.
A hybrid attention mechanism for weakly-supervised temporal action localization
Ashraful Islam, Chengjiang Long, and Richard Radke · 2021
Cited alongside, same era.
Weakly-supervised temporal action localization by uncertainty modeling
Pilhyeon Lee, Jinglu Wang, Yan Lu, and Hyeran Byun · 2021
Cited alongside, same era.
Uncertainty guided collaborative training for weakly supervised temporal action detection
Wenfei Yang, Tianzhu Zhang, Xiaoyuan Yu, Tian Qi, Yongdong Zhang, and Feng Wu · 2021
Cited alongside, same era.
Cola: Weakly-supervised temporal action localization with snippet contrastive learning
Can Zhang, Meng Cao, Dongming Yang, Jie Chen, and Yuexian Zou · 2021
Cited alongside, same era.
Distilling vision-language pre-training to collaborate with weakly-supervised temporal action localization
Chen Ju, Kunhao Zheng, Jinxiang Liu, Peisen Zhao, Ya Zhang, Jianlong Chang, Qi Tian, and Yanfeng Wang · 2023
Later among the works it cites.
GPT-4 Technical Report
OpenAI · 2023
Later among the works it cites.
Proposal-based multiple instance learning for weakly-supervised temporal action localization
Huan Ren, Wenfei Yang, Tianzhu Zhang, and Yongdong Zhang · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Chatvideo: A tracklet-centric multimodal and versatile video understanding system
Junke Wang, Dongdong Chen, Chong Luo, Xiyang Dai, Lu Yuan, Zuxuan Wu, and Yu-Gang Jiang · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Cited alongside, same era.
Dual-evidential learning for weakly-supervised temporal action localization
Mengyuan Chen, Junyu Gao, Shicai Yang, and Changsheng Xu · 2022
Cited alongside, same era.
Exploring denoised cross-video contrast for weakly-supervised temporal action localization
Jingjing Li, Tianyu Yang, Wei Ji, Jue Wang, and Li Cheng · 2022
Cited alongside, same era.
Dynamic graph modeling for weakly-supervised temporal action localization
Haichao Shi, Xiao-Yu Zhang, Changsheng Li, Lixing Gong, Yong Li, and Yongjun Bao · 2022
Cited alongside, same era.
Equivalent classification mapping for weakly supervised temporal action localization
Tao Zhao, Junwei Han, Le Yang, and Dingwen Zhang · 2022
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al · 2023
Cited alongside, same era.
Imagebind: One embedding space to bind them all
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra · 2023
Cited alongside, same era.
Later among the works it cites.
Weakly-supervised temporal action localization by inferring salient snippet-feature
Wulian Yun, Mengshi Qi, Chuanming Wang, and Huadong Ma · 2023
Later among the works it cites.
Video-llama: An instruction-tuned audio-visual language model for video understanding
Hang Zhang, Xin Li, and Lidong Bing · 2023
Later among the works it cites.
Improving weakly supervised temporal action localization by bridging train-test gap in pseudo labels
Jingqiu Zhou, Linjiang Huang, Liang Wang, Si Liu, and Hongsheng Li · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny · 2023
Later among the works it cites.
Gim: A million-scale benchmark for generative image manipulation detection and localization
Yirui Chen, Xudong Huang, Quan Zhang, Wei Li, Mingjian Zhu, Qiangyu Yan, Simiao Li, Hanting Chen, Hailin Hu, Jie Yang, et al · 2024
Closest in time.
Distilling semantic priors from sam to efficient image restoration models
Quan Zhang, Xiaoyu Liu, Wei Li, Hanting Chen, Junchao Liu, Jie Hu, Zhiwei Xiong, Chun Yuan, and Yunhe Wang · 2024
Closest in time.
Modeling multi-task model merging as adaptive projective gradient descent
Yongxian Wei, Anke Tang, Li Shen, Feng Xiong, Chun Yuan, and Xiaochun Cao · 2025
Closest in time.