Fetching the paper…
Reading the bibliography…
Query-based video grounding is an important yet challenging task in video understanding, which aims to localize the target segment in an untrimmed video according to a sentence query.
Temporal sequence modeling for video event detection
Yu Cheng, Quanfu Fan, Sharath Pankanti, and Alok Choudhary · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir D. Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Temporal action localization in untrimmed videos via multi-stage cnns
Zheng Shou, Dongang Wang, and Shih-Fu Chang · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Earlier work this paper cites.
Tall: Temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia · 2017
Earlier work this paper cites.
Tall: Temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Localizing moments in video with temporal language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2018
Earlier work this paper cites.
Recurrent fusion network for image captioning
Wenhao Jiang, Lin Ma, Yu-Gang Jiang, Wei Liu, and Tong Zhang · 2018
Earlier work this paper cites.
Localizing natural language in videos
Jingyuan Chen, Lin Ma, Xinpeng Chen, Zequn Jie, and Jiebo Luo · 2019
Earlier work this paper cites.
Wslln: Weakly supervised natural language localization networks
Mingfei Gao, Larry Davis, Richard Socher, and Caiming Xiong · 2019
Earlier work this paper cites.
Weakly supervised video moment retrieval from text queries
Niluthpol Chowdhury Mithun, Sujoy Paul, and Amit K Roy-Chowdhury · 2019
Earlier work this paper cites.
Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment
Da Zhang, Xiyang Dai, Xin Wang, Yuan-Fang Wang, and Larry S Davis · 2019
Earlier work this paper cites.
Cross-modal interaction networks for query-based moment retrieval in videos
Zhu Zhang, Zhijie Lin, Zhou Zhao, and Zhenxin Xiao · 2019
Cited alongside, same era.
Rethinking the bottom-up framework for query-based video localization
Long Chen, Chujie Lu, Siliang Tang, Jun Xiao, Dong Zhang, Chilie Tan, and Xiaolin Li · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Cited alongside, same era.
Look closer to ground better: Weakly-supervised temporal grounding of sentence in video
Zhenfang Chen, Lin Ma, Wenhan Luo, Peng Tang, and Kwan-Yee K Wong · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Cited alongside, same era.
Dense regression network for video grounding
Runhao Zeng, Haoming Xu, Wenbing Huang, Peihao Chen, Mingkui Tan, and Chuang Gan · 2020
Later among the works it cites.
Span-based localizing network for natural language video localization
Hao Zhang, Aixin Sun, Wei Jing, and Joey Tianyi Zhou · 2020
Later among the works it cites.
Learning 2d temporal adjacent networks for moment localization with natural language
Songyang Zhang, Houwen Peng, Jianlong Fu, and Jiebo Luo · 2020
Later among the works it cites.
Counterfactual contrastive learning for weakly-supervised vision-language grounding
Zhu Zhang, Zhou Zhao, Zhijie Lin, Xiuqiang He, et al · 2020
Later among the works it cites.
Adaptive proposal generation network for temporal sentence localization in videos
Daizong Liu, Xiaoye Qu, Jianfeng Dong, and Pan Zhou · 2021
Later among the works it cites.
Context-aware biaffine localizing network for temporal sentence grounding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Weakly-supervised video moment retrieval via semantic completion network
Zhijie Lin, Zhou Zhao, Zhu Zhang, Qi Wang, and Huasheng Liu · 2020
Cited alongside, same era.
Reasoning step-by-step: Temporal sentence localization in videos via deep rectification-modulation network
Daizong Liu, Xiaoye Qu, Jianfeng Dong, and Pan Zhou · 2020
Cited alongside, same era.
Jointly cross-and self-modal graph attention network for query-based moment localization
Daizong Liu, Xiaoye Qu, Xiao-Yang Liu, Jianfeng Dong, Pan Zhou, and Zichuan Xu · 2020
Cited alongside, same era.
Violin: A large-scale dataset for video-and-language inference
Jingzhou Liu, Wenhu Chen, Yu Cheng, Zhe Gan, Licheng Yu, Yiming Yang, and Jingjing Liu · 2020
Cited alongside, same era.
Vlanet: Video-language alignment network for weakly-supervised video moment retrieval
Minuk Ma, Sunjae Yoon, Junyeong Kim, Youngjoon Lee, Sunghun Kang, and Chang D Yoo · 2020
Cited alongside, same era.
Local-global video-text interactions for temporal grounding
Jonghwan Mun, Minsu Cho, and Bohyung Han · 2020
Cited alongside, same era.
Yijun Song, Jingwen Wang, Lin Ma, Zhou Yu, and Jun Yu · 2020
Cited alongside, same era.
Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou, Yu Cheng, Wei Wei, Zichuan Xu, and Yulai Xie · 2021
Later among the works it cites.
Progressively guide to attend: An iterative alignment framework for temporal sentence grounding
Daizong Liu, Xiaoye Qu, and Pan Zhou · 2021
Later among the works it cites.
Interventional video grounding with dual contrastive learning
Guoshun Nan, Rui Qiao, Yao Xiao, Jun Liu, Sicong Leng, Hao Zhang, and Wei Lu · 2021
Later among the works it cites.
Logan: Latent graph co-attention network for weakly-supervised video moment retrieval
Reuben Tan, Huijuan Xu, Kate Saenko, and Bryan A Plummer · 2021
Later among the works it cites.
Multi-modal relational graph for cross-modal video moment retrieval
Yawen Zeng, Da Cao, Xiaochi Wei, Meng Liu, Zhou Zhao, and Zheng Qin · 2021
Later among the works it cites.
Video corpus moment retrieval with contrastive learning
Hao Zhang, Aixin Sun, Wei Jing, Guoshun Nan, Liangli Zhen, Joey Tianyi Zhou, and Rick Siow Mong Goh · 2021
Later among the works it cites.
Exploring optical-flow-guided motion and detection-based appearance for temporal sentence grounding
Daizong Liu, Xiang Fang, Wei Hu, and Pan Zhou · 2022
Closest in time.
Memory-guided semantic learning network for temporal sentence grounding
Daizong Liu, Xiaoye Qu, Xing Di, Yu Cheng, Zichuan Xu Xu, and Pan Zhou · 2022
Closest in time.
Unsupervised temporal video grounding with deep semantic clustering
Daizong Liu, Xiaoye Qu, Yinzhen Wang, Xing Di, Kai Zou, Yu Cheng, Zichuan Xu, and Pan Zhou · 2022
Closest in time.
Exploring motion and appearance information for temporal sentence grounding
Daizong Liu, Xiaoye Qu, Pan Zhou, and Yang Liu · 2022
Closest in time.