Fetching the paper…
Reading the bibliography…
Video Moment Retrieval and Highlight Detection aim to find corresponding content in the video based on a text query.
Learning representations by back-propagating errors
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams · 1986
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Grounding action descriptions in videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Multi-task deep visual-semantic embedding for video thumbnail selection
Wu Liu, Tao Mei, Yongdong Zhang, Cherry Che, and Jiebo Luo · 2015
Earlier work this paper cites.
Tvsum: Summarizing web videos using titles
Yale Song, Jordi Vallmitjana, Amanda Stent, and Alejandro Jaimes · 2015
Earlier work this paper cites.
Video2gif: Automatic generation of animated gifs from video
Michael Gygli, Yale Song, and Liangliang Cao · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
To click or not to click: Automatic selection of beautiful thumbnails from videos
Yale Song, Miriam Redi, Jordi Vallmitjana, and Alejandro Jaimes · 2016
Earlier work this paper cites.
Highlight detection with pairwise deep ranking for first-person video summarization
Ting Yao, Tao Mei, and Yong Rui · 2016
Earlier work this paper cites.
Video summarization with long short-term memory
Ke Zhang, Wei-Lun Chao, Fei Sha, and Kristen Grauman · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Earlier work this paper cites.
Tall: Temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I Loshchilov · 2017
Earlier work this paper cites.
Unsupervised video summarization with adversarial lstm networks
Behrooz Mahasseni, Michael Lam, and Sinisa Todorovic · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Weakly-supervised video summarization using variational encoder-decoder and web prior
Sijia Cai, Wangmeng Zuo, Larry S Davis, and Lei Zhang · 2018
Earlier work this paper cites.
Temporally grounding natural sentence in video
Jingyuan Chen, Xinpeng Chen, Lin Ma, Zequn Jie, and Tat-Seng Chua · 2018
Earlier work this paper cites.
Cross-modal moment localization in videos
Meng Liu, Xiang Wang, Liqiang Nie, Qi Tian, Baoquan Chen, and Tat-Seng Chua · 2018
Earlier work this paper cites.
A deep ranking model for spatio-temporal highlight detection from a 360◦ video
Youngjae Yu, Sangho Lee, Joonil Na, Jaeyun Kang, and Gunhee Kim · 2018
Earlier work this paper cites.
Semantic proposal for activity localization in videos via sentence query
Shaoxiang Chen and Yu-Gang Jiang · 2019
Earlier work this paper cites.
Finding moments in video collections using natural language
Victor Escorcia, Mattia Soldan, Josef Sivic, Bernard Ghanem, and Bryan Russell · 2019
Earlier work this paper cites.
Temporal localization of moments in video collections with natural language
Victor Escorcia, Mattia Soldan, Josef Sivic, Bernard Ghanem, and Bryan Russell · 2019
Earlier work this paper cites.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Earlier work this paper cites.
Generalized intersection over union: A metric and a loss for bounding box regression
Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese · 2019
Earlier work this paper cites.
Language-driven temporal activity localization: A semantic matching reinforcement learning model
Weining Wang, Yan Huang, and Liang Wang · 2019
Earlier work this paper cites.
Less is more: Learning highlight detection from video duration
Bo Xiong, Yannis Kalantidis, Deepti Ghadiyaram, and Kristen Grauman · 2019
Earlier work this paper cites.
Unsupervised video summarization with cycle-consistent adversarial lstm networks
Li Yuan, Francis Eng Hock Tay, Ping Li, and Jiashi Feng · 2019
Earlier work this paper cites.
Semantic conditioned dynamic modulation for temporal sentence grounding in videos
Yitian Yuan, Lin Ma, Jingwen Wang, Wei Liu, and Wenwu Zhu · 2019
Earlier work this paper cites.
To find where you talk: Temporal sentence localization in video with attention based location regression
Yitian Yuan, Tao Mei, and Wenwu Zhu · 2019
Earlier work this paper cites.
Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment
Da Zhang, Xiyang Dai, Xin Wang, Yuan-Fang Wang, and Larry S Davis · 2019
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Earlier work this paper cites.
Cross-modal video moment retrieval based on visual-textual relationship alignment
Zhuo Chen, Hao Du, Yufei Wu, Tong Xu, and Enhong Chen · 2020
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Earlier work this paper cites.
Mini-net: Multiple instance ranking network for video highlight detection
Fa-Ting Hong, Xuanteng Huang, Wei-Hong Li, and Wei-Shi Zheng · 2020
Cited alongside, same era.
Panns: Large-scale pretrained audio neural networks for audio pattern recognition
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D Plumbley · 2020
Cited alongside, same era.
Tvr: A large-scale dataset for video-subtitle moment retrieval
Jie Lei, Licheng Yu, Tamara L Berg, and Mohit Bansal · 2020
Cited alongside, same era.
Contrastive multiview coding
Yonglong Tian, Dilip Krishnan, and Phillip Isola · 2020
Cited alongside, same era.
Learning trailer moments in full-length movies with co-contrastive attention
Lezi Wang, Dong Liu, Rohit Puri, and Dimitris N Metaxas · 2020
Cited alongside, same era.
Dense regression network for video grounding
Runhao Zeng, Haoming Xu, Wenbing Huang, Peihao Chen, Mingkui Tan, and Chuang Gan · 2020
Cross-modal video moment retrieval based on enhancing significant features
Jinfu Yang, Yubin Liu, Lin Song, and Xue Yan · 2022
Later among the works it cites.
Temporal sentence grounding in videos with fine-grained multimodal correlation
Yitian Yuan, Xin Wang, and Wenwu Zhu · 2022
Later among the works it cites.
Momentum cross-modal contrastive learning for video moment retrieval
De Han, Xing Cheng, Nan Guo, Xiaochun Ye, Benjamin Rainer, and Peter Priller · 2023
Later among the works it cites.
Knowing where to focus: Event-aware transformer for video grounding
Jinhyun Jang, Jungin Park, Jin Kim, Hyeongjun Kwon, and Kwanghoon Sohn · 2023
Later among the works it cites.
Are binary annotations sufficient? video moment retrieval via hierarchical uncertainty-based active learning
Wei Ji, Renjie Liang, Zhedong Zheng, Wenqiao Zhang, Shengyu Zhang, Juncheng Li, Mengze Li, and Tat-seng Chua · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Span-based localizing network for natural language video localization
Hao Zhang, Aixin Sun, Wei Jing, and Joey Tianyi Zhou · 2020
Cited alongside, same era.
Learning 2d temporal adjacent networks for moment localization with natural language
Songyang Zhang, Houwen Peng, Jianlong Fu, and Jiebo Luo · 2020
Cited alongside, same era.
End-to-end object detection with adaptive clustering transformer
Minghang Zheng, Peng Gao, Renrui Zhang, Kunchang Li, Xiaogang Wang, Hongsheng Li, and Hao Dong · 2020
Cited alongside, same era.
Deformable detr: Deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai · 2020
Cited alongside, same era.
Joint visual and audio learning for video highlight detection
Taivanbat Badamdorj, Mrigank Rochan, Yang Wang, and Li Cheng · 2021
Cited alongside, same era.
Exploring simple siamese representation learning
Xinlei Chen and Kaiming He · 2021
Cited alongside, same era.
Efficient weakly-supervised video moment retrieval without multimodal fusion
Xun Jiang, Xing Xu, Fumin Shen, Guoqing Wang, and Yang Yang · 2023
Later among the works it cites.
MS-DETR: Natural language video localization with sampling moment-moment interaction
Wang Jing, Aixin Sun, Hao Zhang, and Xiaoli Li · 2023
Later among the works it cites.
Overcoming weak visual-textual alignment for video moment retrieval
Minjoon Jung, Youwon Jang, Seongho Choi, Joochan Kim, Jin-Hwa Kim, and Byoung-Tak Zhang · 2023
Later among the works it cites.
Bam-detr: Boundary-aligned moment detection transformer for temporal sentence grounding in videos
Pilhyeon Lee and Hyeran Byun · 2023
Later among the works it cites.
Mim: Lightweight multi-modal interaction model for joint video moment retrieval and highlight detection
Jinyu Li, Fuwei Zhang, Shujin Lin, Fan Zhou, and Ruomei Wang · 2023
Later among the works it cites.
Univtg: Towards unified video-language temporal grounding
Kevin Qinghong Lin, Pengchuan Zhang, Joya Chen, Shraman Pramanick, Difei Gao, Alex Jinpeng Wang, Rui Yan, and Mike Zheng Shou · 2023
Later among the works it cites.
Llavilo: Boosting video moment retrieval via adapter-based multimodal modeling
Kaijing Ma, Xianghao Zang, Zerun Feng, Han Fang, Chao Ban, Yuhan Wei, Zhongjiang He, Yongxiang Li, and Hao Sun · 2023
Later among the works it cites.
WonJun Moon, Sangeek Hyun, SuBeen Lee, and Jae-Pil Heo · 2023
Later among the works it cites.
Query-dependent video representation for moment retrieval and highlight detection
WonJun Moon, Sangeek Hyun, SangUk Park, Dongchan Park, and Jae-Pil Heo · 2023
Later among the works it cites.
Learning grounded vision-language representation for versatile understanding in untrimmed videos
Teng Wang, Jinrui Zhang, Feng Zheng, Wenhao Jiang, Ran Cheng, and Ping Luo · 2023
Later among the works it cites.
Unloc: A unified framework for video localization tasks
Shen Yan, Xuehan Xiong, Arsha Nagrani, Anurag Arnab, Zhonghao Wang, Weina Ge, David Ross, and Cordelia Schmid · 2023
Later among the works it cites.
Temporally language grounding with multi-modal multi-prompt tuning
Yawen Zeng, Ning Han, Keyu Pan, and Qin Jin · 2023
Later among the works it cites.
Video mamba suite: State space model as a versatile alternative for video understanding
Guo Chen, Yifei Huang, Jilan Xu, Baoqi Pei, Zhe Chen, Zhiqi Li, Jiahao Wang, Kunchang Li, Tong Lu, and Limin Wang · 2024
Later among the works it cites.
Semantic fusion augmentation and semantic boundary detection: A novel approach to multi-target video moment retrieval
Cheng Huang, Yi-Lun Wu, Hong-Han Shuai, and Ching-Chun Huang · 2024
Later among the works it cites.
Prior knowledge integration via llm encoding and pseudo event regulation for video moment retrieval
Yiyang Jiang, Wengyu Zhang, Xulu Zhang, Xiaoyong Wei, Chang Wen Chen, and Qing Li · 2024
Later among the works it cites.
Research on video content grounding based on cross-modal information interaction
Jinyu Li · 2024
Later among the works it cites.
Momentdiff: Generative video moment retrieval from random to real
Pandeng Li, Chen-Wei Xie, Hongtao Xie, Liming Zhao, Lei Zhang, Yun Zheng, Deli Zhao, and Yongdong Zhang · 2024
Later among the works it cites.
Context-enhanced video moment retrieval with large language models
Weijia Liu, Bo Miao, Jiuxin Cao, Xuelin Zhu, Bo Liu, Mehwish Nasim, and Ajmal Mian · 2024
Later among the works it cites.
r 2 r^{2} -tuning: Efficient image-to-video transfer learning for video temporal grounding
Ye Liu, Jixuan He, Wanhua Li, Junsik Kim, Donglai Wei, Hanspeter Pfister, and Chang Wen Chen · 2024
Later among the works it cites.
Towards balanced alignment: Modal-enhanced semantic modeling for video moment retrieval
Zhihang Liu, Jun Li, Hongtao Xie, Pandeng Li, Jiannan Ge, Sun-Ao Liu, and Guoqing Jin · 2024
Later among the works it cites.
Disentangle and denoise: Tackling context misalignment for video moment retrieval
Kaijing Ma, Han Fang, Xianghao Zang, Chao Ban, Lanxiang Zhou, Zhongjiang He, Yongxiang Li, Hao Sun, Zerun Feng, and Xingsong Hou · 2024
Later among the works it cites.
Pre-training model for video moment retrieval based on clip
Li Miao, Weifen Zhang, and Ling Xu · 2024
Later among the works it cites.
Cross-modal contrastive learning with asymmetric co-attention network for video moment retrieval
Love Panta, Prashant Shrestha, Brabeem Sapkota, Amrita Bhattarai, Suresh Manandhar, and Anand Kumar Sah · 2024
Later among the works it cites.
Tr-detr: Task-reciprocal transformer for joint moment retrieval and highlight detection
Hao Sun, Mingyao Zhou, Wenjing Chen, and Wei Xie · 2024
Later among the works it cites.
Modality-aware heterogeneous graph for joint video moment retrieval and highlight detection
Ruomei Wang, Jiawei Feng, Fuwei Zhang, Xiaonan Luo, and Yuanmao Luo · 2024
Later among the works it cites.
Bridging the gap: A unified video comprehension framework for moment retrieval and highlight detection
Yicheng Xiao, Zhuoyan Luo, Yong Liu, Yue Ma, Hengwei Bian, Yatai Ji, Yujiu Yang, and Xiu Li · 2024
Later among the works it cites.
Mh-detr: Video moment and highlight detection with cross-modal transformer
Yifang Xu, Yunzhuo Sun, Benxiang Zhai, Youyao Jia, and Sidan Du · 2024
Later among the works it cites.
Task-driven exploration: Decoupling and inter-task feedback for joint moment retrieval and highlight detection
Jin Yang, Ping Wei, Huan Li, and Ziyang Ren · 2024
Later among the works it cites.
Self-chained image-language model for video localization and question answering
Shoubin Yu, Jaemin Cho, Prateek Yadav, and Mohit Bansal · 2024
Later among the works it cites.
Subtask prior-driven optimized mechanism on joint video moment retrieval and highlight detection
Siyu Zhou, Fuwei Zhang, Ruomei Wang, Fan Zhou, and Zhuo Su · 2024
Later among the works it cites.