Fetching the paper…
Reading the bibliography…
In this paper, we investigate the feasibility of leveraging large language models (LLMs) for integrating general knowledge and incorporating pseudo-events as priors for temporal content distribution in video moment retrieval (VMR) models.
Less is More: Learning Highlight Detection from Video Duration
Bo Xiong, Yannis Kalantidis, Deepti Ghadiyaram, and Kristen Grauman. 2019b · 1903
Earlier work this paper cites.
TVR: A Large-Scale Dataset for Video-Subtitle Moment Retrieval
Jie Lei, Licheng Yu, Tamara L. Berg, and Mohit Bansal. 2020 · 2001
Earlier work this paper cites.
End-to-End Object Detection with Transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020 · 2005
Earlier work this paper cites.
MINI-Net: Multiple Instance Ranking Network for Video Highlight Detection
Fa-Ting Hong, Xuanteng Huang, Wei-Hong Li, and Wei-Shi Zheng. 2020 · 2007
Earlier work this paper cites.
Learning Trailer Moments in Full-Length Movies
Lezi Wang, Dong Liu, Rohit Puri, and Dimitris N. Metaxas. 2020 · 2008
Earlier work this paper cites.
Fusing semantics, observability, reliability and diversity of concept detectors for video search. In Proceedings of the 16th ACM International Conference on Multimedia (Vancouver, British Columbia, Canada) (MM ’08) . Association for Computing Machinery, New York, NY, USA, 81–90
Xiao-Yong Wei and Chong-Wah Ngo. 2008 · 2008
Earlier work this paper cites.
Coached active learning for interactive video search. In Proceedings of the 19th ACM International Conference on Multimedia (Scottsdale, Arizona, USA) (MM ’11) . Association for Computing Machinery, New York, NY, USA, 443–452
Xiao-Yong Wei and Zhen-Qun Yang. 2011 · 2011
Earlier work this paper cites.
Coaching the Exploration and Exploitation in Active Learning for Interactive Video Retrieval
Xiao-Yong Wei and Zhen-Qun Yang. 2013 · 2012
Earlier work this paper cites.
Grounding Action Descriptions in Videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal. 2013 · 2013
Earlier work this paper cites.
Coherent Multi-sentence Video Description with Variable Level of Detail
Anna Rohrbach, Marcus Rohrbach, Wei Qiu, Annemarie Friedrich, Manfred Pinkal, and Bernt Schiele. 2014 · 2014
Earlier work this paper cites.
Ranking Domain-specific Highlights by Analyzing Edited Videos. In ECCV
Min Sun, Ali Farhadi, and Steve Seitz. 2014 · 2014
Earlier work this paper cites.
Multi-task deep visual-semantic embedding for video thumbnail selection. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 3707–3715
Wu Liu, Tao Mei, Yongdong Zhang, Cherry Che, and Jiebo Luo. 2015 · 2015
Earlier work this paper cites.
Very Deep Convolutional Networks for Large-Scale Image Recognition
Karen Simonyan and Andrew Zisserman. 2015 · 2015
Earlier work this paper cites.
TVSum: Summarizing Web Videos Using Titles. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Yale Song, Jordi Vallmitjana, Amanda Stent, and Alejandro Jaimes. 2015 · 2015
Earlier work this paper cites.
Unsupervised Extraction of Video Highlights Via Robust Recurrent Auto-encoders
Huan Yang, Baoyuan Wang, Stephen Lin, David Wipf, Minyi Guo, and Baining Guo. 2015 · 2015
Earlier work this paper cites.
Video2gif: Automatic generation of animated gifs from video. In Proceedings of the IEEE conference on computer vision and pattern recognition . 1001–1009
Michael Gygli, Yale Song, and Liangliang Cao. 2016 · 2016
Earlier work this paper cites.
To Click or Not To Click: Automatic Selection of Beautiful Thumbnails from Videos
Yale Song, Miriam Redi, Jordi Vallmitjana, and Alejandro Jaimes. 2016 · 2016
Earlier work this paper cites.
Video Summarization with Long Short-term Memory
Ke Zhang, Wei-Lun Chao, Fei Sha, and Kristen Grauman. 2016 · 2016
Earlier work this paper cites.
Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 4724–4733
João Carreira and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
Tall: Temporal activity localization via language query. In Proceedings of the IEEE international conference on computer vision . 5267–5275
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia. 2017 · 2017
Earlier work this paper cites.
Localizing moments in video with natural language. In Proceedings of the IEEE international conference on computer vision . 5803–5812
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. 2017 · 2017
Earlier work this paper cites.
Fixing Weight Decay Regularization in Adam
Ilya Loshchilov and Frank Hutter. 2017 · 2017
Earlier work this paper cites.
Unsupervised Video Summarization with Adversarial LSTM Networks. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 2982–2991
Behrooz Mahasseni, Michael Lam, and Sinisa Todorovic. 2017b · 2017
Cited alongside, same era.
Weakly-supervised Video Summarization using Variational Encoder-Decoder and Web Prior. In Proceedings of the European Conference on Computer Vision (ECCV)
Sijia Cai, Wangmeng Zuo, Larry S. Davis, and Lei Zhang. 2018 · 2018
Cited alongside, same era.
Attentive moment retrieval in videos. In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval . 15–24
Meng Liu, Xiang Wang, Liqiang Nie, Xiangnan He, Baoquan Chen, and Tat-Seng Chua. 2018 · 2018
Cited alongside, same era.
Temporal Localization of Moments in Video Collections with Natural Language
Victor Escorcia, Mattia Soldan, Josef Sivic, Bernard Ghanem, and Bryan Russell. 2019 · 2019
Cited alongside, same era.
SlowFast Networks for Video Recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Later among the works it cites.
GPT-4V(ision) System Card
2023 · 2023
Later among the works it cites.
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023 · 2023
Later among the works it cites.
MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Jun Chen, Deyao Zhu, Xiaoqian Shen, Xiang Li, Zechun Liu, Pengchuan Zhang, Raghuraman Krishnamoorthi, Vikas Chandra, Yunyang Xiong, and Mohamed Elhoseiny. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. 2019 · 2019
Cited alongside, same era.
Generalized Intersection Over Union: A Metric and a Loss for Bounding Box Regression. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 658–666
Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese. 2019 · 2019
Cited alongside, same era.
Semantic conditioned dynamic modulation for temporal sentence grounding in videos
Yitian Yuan, Lin Ma, Jingwen Wang, Wei Liu, and Wenwu Zhu. 2019 · 2019
Cited alongside, same era.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Cited alongside, same era.
Adaptive video highlight detection by learning from user history. In European conference on computer vision . Springer, 261–278
Mrigank Rochan, Mahesh Kumar Krishna Reddy, Linwei Ye, and Yang Wang. 2020 · 2020
Cited alongside, same era.
Joint Visual and Audio Learning for Video Highlight Detection. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV) . 8107–8117
Taivanbat Badamdorj, Mrigank Rochan, Yang Wang, and Li Cheng. 2021 · 2021
Cited alongside, same era.
QVHighlights: Detecting Moments and Highlights in Videos via Natural Language Queries
Jie Lei, Tamara L. Berg, and Mohit Bansal. 2021 · 2021
Cited alongside, same era.
Context-aware biaffine localizing network for temporal sentence grounding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 11235–11244
Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou, Yu Cheng, Wei Wei, Zichuan Xu, and Yulai Xie. 2021 · 2021
Cited alongside, same era.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023 · 2023
Later among the works it cites.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. 2023 · 2023
Later among the works it cites.
Beyond Visual Cues: Synchronously Exploring Target-Centric Semantics for Vision-Language Tracking
Jiawei Ge, Xiangmei Chen, Jiuxin Cao, Xuelin Zhu, Weijia Liu, and Bo Liu. 2023 · 2023
Later among the works it cites.
Knowing Where to Focus: Event-aware Transformer for Video Grounding. In Proceedings of the IEEE/CVF International Conference on Computer Vision
Jinhyun Jang, Jungin Park, Jin Kim, Hyeongjun Kwon, and Kwanghoon Sohn. 2023 · 2023
Later among the works it cites.
UniVTG: Towards Unified Video-Language Temporal Grounding
Kevin Qinghong Lin, Pengchuan Zhang, Joya Chen, Shraman Pramanick, Difei Gao, Alex Jinpeng Wang, Rui Yan, and Mike Zheng Shou. 2023 · 2023
Later among the works it cites.
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023 · 2023
Later among the works it cites.
LLaViLo: Boosting Video Moment Retrieval via Adapter-Based Multimodal Modeling. In 2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW) . 2790–2795
Kaijing Ma, Xianghao Zang, Zerun Feng, Han Fang, Chao Ban, Yuhan Wei, Zhongjiang He, Yongxiang Li, and Hao Sun. 2023 · 2023
Later among the works it cites.
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. 2023 · 2023
Later among the works it cites.
WonJun Moon, Sangeek Hyun, SuBeen Lee, and Jae-Pil Heo. 2023a · 2023
Later among the works it cites.
Frozen Transformers in Language Models Are Effective Visual Encoder Layers
Ziqi Pang, Ziyang Xie, Yunze Man, and Yu-Xiong Wang. 2023 · 2023
Later among the works it cites.
Kosmos-2: Grounding Multimodal Large Language Models to the World
Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. 2023 · 2023
Later among the works it cites.
Generative Pretraining in Multimodality
Quan Sun, Qiying Yu, Yufeng Cui, Fan Zhang, Xiaosong Zhang, Yueze Wang, Hongcheng Gao, Jingjing Liu, Tiejun Huang, and Xinlong Wang. 2023 · 2023
Later among the works it cites.
InternLM: A Multilingual Language Model with Progressively Enhanced Capabilities
InternLM Team. 2023 · 2023
Later among the works it cites.
Unloc: A unified framework for video localization tasks. In Proceedings of the IEEE/CVF International Conference on Computer Vision
Shen Yan, Xuehan Xiong, Arsha Nagrani, Anurag Arnab, Zhonghao Wang, Weina Ge, David Ross, and Cordelia Schmid. 2023 · 2023
Later among the works it cites.
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023 · 2023
Later among the works it cites.
VideoChat: Chat-Centric Video Understanding
KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. 2024 · 2024
Closest in time.
Gemini: A Family of Highly Capable Multimodal Models
Gemini Team. 2024 · 2024
Closest in time.
MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens
Kaizhi Zheng, Xuehai He, and Xin Eric Wang. 2024 · 2024
Closest in time.