Fetching the paper…
Reading the bibliography…
Temporal sentence grounding in videos (TSGV), \aka natural language video localization (NLVL) or video moment retrieval (VMR), aims to retrieve a temporal moment that semantically corresponds to a language query from an untrimmed video.
A. Shapiro, “Monte carlo sampling methods,” Handbooks in operations research and management science , vol. 10, 2003
2003
Earlier work this paper cites.
O. Pele and M. Werman, “Fast and robust earth mover’s distances,” in ICCV , 2009
2009
Earlier work this paper cites.
M. Rohrbach, M. Regneri, M. Andriluka, S. Amin, M. Pinkal, and B. Schiele, “Script data for attribute-based recognition of composite activities,” in ECCV , 2012
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
M. Regneri, M. Rohrbach, D. Wetzel, S. Thater, B. Schiele, and M. Pinkal, “Grounding action descriptions in videos,” TACL , vol. 1, 2013
2013
Earlier work this paper cites.
J. Pennington, R. Socher, and C. Manning, “GloVe: Global vectors for word representation,” in EMNLP , 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
C. Manning, M. Surdeanu, J. Bauer, J. Finkel, S. Bethard, and D. McClosky, “The Stanford CoreNLP natural language processing toolkit,” in ACL: System Demonstrations , 2014
2014
Earlier work this paper cites.
R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in CVPR , 2014
2014
Earlier work this paper cites.
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. B. Girshick, S. Guadarrama, and T. Darrell, “Caffe: Convolutional architecture for fast feature embedding,” ACM MM , 2014
2014
Earlier work this paper cites.
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learning spatiotemporal features with 3d convolutional networks,” in ICCV , 2015
2015
Earlier work this paper cites.
R. Kiros, Y. Zhu, R. R. Salakhutdinov, R. Zemel, R. Urtasun, A. Torralba, and S. Fidler, “Skip-thought vectors,” in NeurIPS , vol. 28, 2015
2015
Earlier work this paper cites.
F. C. Heilbron, V. Escorcia, B. Ghanem, and J. C. Niebles, “Activitynet: A large-scale video benchmark for human activity understanding,” in CVPR , 2015
2015
Earlier work this paper cites.
C. Feichtenhofer, A. Pinz, and A. Zisserman, “Convolutional two-stream network fusion for video action recognition,” in CVPR , 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016
2016
Earlier work this paper cites.
B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li, “Yfcc100m: The new data in multimedia research,” Communications of the ACM , vol. 59, 2016
2016
Earlier work this paper cites.
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta, “Hollywood in homes: Crowdsourcing data collection for activity understanding,” in ECCV , 2016
2016
Earlier work this paper cites.
R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell, “Natural language object retrieval,” in CVPR , 2016
2016
Earlier work this paper cites.
Y. Aytar, C. Vondrick, and A. Torralba, “Soundnet: Learning sound representations from unlabeled video,” in NeurIPS , vol. 29, 2016
2016
Earlier work this paper cites.
X. Ma and E. Hovy, “End-to-end sequence labeling via bi-directional lstm-cnns-crf,” in ACL , 2016
2016
Earlier work this paper cites.
J. Pearl, M. Glymour, and N. P. Jewell, Causal inference in statistics: A primer . John Wiley & Sons, 2016
2016
Earlier work this paper cites.
Y. Pan, T. Mei, T. Yao, H. Li, and Y. Rui, “Jointly modeling embedding and translation to bridge video and language,” in CVPR , 2016
2016
Earlier work this paper cites.
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in CVPR , 2017
2017
Earlier work this paper cites.
W. Wang, N. Yang, F. Wei, B. Chang, and M. Zhou, “Gated self-matching networks for reading comprehension and question answering,” in ACL , 2017
2017
Earlier work this paper cites.
M. Seo, A. Kembhavi, A. Farhadi, and H. Hajishirzi, “Bidirectional attention flow for machine comprehension,” in ICLR , 2017
2017
Earlier work this paper cites.
J. Gao, C. Sun, Z. Yang, and R. Nevatia, “Tall: Temporal activity localization via language query,” in ICCV , 2017
2017
Earlier work this paper cites.
L. A. Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell, “Localizing moments in video with natural language,” in ICCV , 2017
2017
Earlier work this paper cites.
A. Conneau, D. Kiela, H. Schwenk, L. Barrault, and A. Bordes, “Supervised learning of universal sentence representations from natural language inference data,” in EMNLP , 2017
2017
Earlier work this paper cites.
H. Xu, A. Das, and K. Saenko, “R-c3d: Region convolutional 3d network for temporal activity detection,” in ICCV , 2017
2017
Earlier work this paper cites.
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. C. Niebles, “Dense-captioning events in videos,” in ICCV , 2017
2017
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” IEEE TPAMI , vol. 39, 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS , vol. 30, 2017
2017
Earlier work this paper cites.
Y. Zhao, Y. Xiong, L. Wang, Z. Wu, X. Tang, and D. Lin, “Temporal action detection with structured segment networks,” in ICCV , 2017
2017
Earlier work this paper cites.
C. Szegedy, S. Ioffe, V. Vanhoucke, and A. A. Alemi, “Inception-v4, inception-resnet and the impact of residual connections on learning,” in AAAI , 2017
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in ICASSP , 2017
2017
Earlier work this paper cites.
Y. Jang, Y. Song, Y. Yu, Y. Kim, and G. Kim, “TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering,” in CVPR , 2017
2017
Earlier work this paper cites.
A. W. Yu, D. Dohan, Q. Le, T. Luong, R. Zhao, and K. Chen, “Fast and accurate reading comprehension by combining self-attention and convolution,” in ICLR , 2018
2018
Earlier work this paper cites.
H. Huang, C. Zhu, Y. Shen, and W. Chen, “Fusionnet: Fusing via fully-aware attention with application to machine comprehension,” in ICLR , 2018
2018
Earlier work this paper cites.
M. Liu, X. Wang, L. Nie, Q. Tian, B. Chen, and T.-S. Chua, “Cross-modal moment localization in videos,” in ACM MM , 2018
2018
Earlier work this paper cites.
M. Liu, X. Wang, L. Nie, X. He, B. Chen, and T.-S. Chua, “Attentive moment retrieval in videos,” in SIGIR , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Chen, X. Chen, L. Ma, Z. Jie, and T.-S. Chua, “Temporally grounding natural sentence in video,” in EMNLP , 2018
2018
Earlier work this paper cites.
X. Duan, W. Huang, C. Gan, J. Wang, W. Zhu, and J. Huang, “Weakly supervised dense event captioning in videos,” in NeurIPS , vol. 31, 2018
2018
Earlier work this paper cites.
L. A. Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell, “Localizing moments in video with temporal language,” in EMNLP , 2018
2018
Earlier work this paper cites.
A. Wu and Y. Han, “Multi-modal circulant fusion for video-to-language and backward,” in IJCAI , 2018
2018
Earlier work this paper cites.
B. Liu, S. Yeung, E. Chou, D.-A. Huang, L. Fei-Fei, and J. C. Niebles, “Temporal modular networks for retrieving complex compositional activities in videos,” in ECCV , 2018
2018
Earlier work this paper cites.
D. Shao, Y. Xiong, Y. Zhao, Q. Huang, Y. Qiao, and D. Lin, “Find and focus: Retrieve and localize video events with natural language queries,” in ECCV , 2018
2018
Earlier work this paper cites.
C. Clark and M. Gardner, “Simple and effective multi-paragraph reading comprehension,” in ACL , 2018
2018
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction . MIT press, 2018
2018
Earlier work this paper cites.
D. Li, H. Wu, J. Zhang, and K. Huang, “A2-rl: Aesthetics aware reinforcement learning for image cropping,” in CVPR , 2018
2018
Earlier work this paper cites.
D.-A. Huang, S. Buch, L. Dery, A. Garg, L. Fei-Fei, and J. C. Niebles, “Finding ”it”: Weakly-supervised reference-aware visual grounding in instructional videos,” in CVPR , 2018
2018
Earlier work this paper cites.
Y. Tian, J. Shi, B. Li, Z. Duan, and C. Xu, “Audio-visual event localization in unconstrained videos,” in ECCV , 2018
2018
Earlier work this paper cites.
N. Garcia and G. Vogiatzis, “Asymmetric spatio-temporal embeddings for large-scale image-to-video retrieval,” in BMVC , 2018
2018
Earlier work this paper cites.
Y. Feng, L. Ma, W. Liu, T. Zhang, and J. Luo, “Video re-localization,” in ECCV , 2018
2018
Earlier work this paper cites.
T. Afouras, J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman, “Deep audio-visual speech recognition,” IEEE TPAMI , 2018
2018
Earlier work this paper cites.
C. Deng, Q. Wu, Q. Wu, F. Hu, F. Lyu, and M. Tan, “Visual grounding via accumulated attention,” in CVPR , 2018
2018
Earlier work this paper cites.
J. Lei, L. Yu, M. Bansal, and T. Berg, “TVQA: Localized, compositional video question answering,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , 2018
2018
Earlier work this paper cites.
J. Gao, R. Ge, K. Chen, and R. Nevatia, “Motion-appearance co-memory networks for video question answering,” in CVPR , 2018
2018
Earlier work this paper cites.
R. Pasunuru and M. Bansal, “Game-based video-context dialogue,” ArXiv , vol. abs/1809.04560, 2018
2018
Earlier work this paper cites.
R. Ge, J. Gao, K. Chen, and R. Nevatia, “Mac: Mining activity concepts for language-based temporal localization,” in WACV , 2019
2019
Earlier work this paper cites.
D. Zhang, X. Dai, X. Wang, Y.-F. Wang, and L. S. Davis, “Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment,” in CVPR , 2019
2019
Earlier work this paper cites.
Y. Yuan, L. Ma, J. Wang, W. Liu, and W. Zhu, “Semantic conditioned dynamic modulation for temporal sentence grounding in videos,” in NeurIPS , 2019
2019
Earlier work this paper cites.
Y. Yuan, T. Mei, and W. Zhu, “To find where you talk: Temporal sentence localization in video with attention based location regression,” in AAAI , vol. 33, 2019
2019
Earlier work this paper cites.
S. Ghosh, A. Agarwal, Z. Parekh, and A. Hauptmann, “ExCL: Extractive Clip Localization Using Natural Language Descriptions,” in NAACL , 2019
2019
Earlier work this paper cites.
J. Chen, L. Ma, X. Chen, Z. Jie, and J. Luo, “Localizing natural language in videos,” in AAAI , vol. 33, 2019
2019
Earlier work this paper cites.
C. Lu, L. Chen, C. Tan, X. Li, and J. Xiao, “DEBUG: A dense bottom-up grounding approach for natural language video localization,” in EMNLP , 2019
2019
Earlier work this paper cites.
D. He, X. Zhao, J. Huang, F. Li, X. Liu, and S. Wen, “Read, watch, and move: Reinforcement learning for temporally grounding natural language descriptions in videos,” in AAAI , vol. 33, 2019
2019
Earlier work this paper cites.
W. Wang, Y. Huang, and L. Wang, “Language-driven temporal activity localization: A semantic matching reinforcement learning model,” in CVPR , 2019
2019
Earlier work this paper cites.
N. C. Mithun, S. Paul, and A. K. Roy-Chowdhury, “Weakly supervised video moment retrieval from text queries,” in CVPR , 2019
2019
Earlier work this paper cites.
M. Gao, L. Davis, R. Socher, and C. Xiong, “WSLLN:weakly supervised natural language localization networks,” in EMNLP , 2019
2019
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in NAACL , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using Siamese BERT-networks,” in EMNLP , 2019
2019
Earlier work this paper cites.
S. Zhang, J. Su, and J. Luo, “Exploiting temporal relationships in video moment localization with natural language,” in ACM MM , 2019
2019
Earlier work this paper cites.
B. Jiang, X. Huang, C. Yang, and J. Yuan, “Cross-modal video moment retrieval with spatial and language-temporal attention,” in ACM ICMR , 2019
2019
Earlier work this paper cites.
H. Xu, K. He, B. A. Plummer, L. Sigal, S. Sclaroff, and K. Saenko, “Multilevel language and vision integration for text-to-clip retrieval,” in AAAI , vol. 33, 2019
2019
Earlier work this paper cites.
S. Chen and Y.-G. Jiang, “Semantic proposal for activity localization in videos via sentence query,” in AAAI , vol. 33, 2019
2019
Earlier work this paper cites.
Z. Zhang, Z. Lin, Z. Zhao, and Z. Xiao, “Cross-modal interaction networks for query-based moment retrieval in videos,” in SIGIR , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. T. Zhou, H. Zhang, D. Jin, H. Zhu, M. Fang, R. S. M. Goh, and K. Kwok, “Dual adversarial neural transfer for low-resource named entity recognition,” in ACL , 2019
2019
Earlier work this paper cites.
T. Lin, X. Liu, X. Li, E. Ding, and S. Wen, “Bmn: Boundary-matching network for temporal action proposal generation,” in ICCV , 2019
2019
Earlier work this paper cites.
J. Shi, J. Xu, B. Gong, and C. Xu, “Not all frames are equal: Weakly-supervised video grounding with contextual similarity and visual clustering losses,” in CVPR , 2019
2019
Earlier work this paper cites.
Z. Chen, L. Ma, W. Luo, and K.-Y. K. Wong, “Weakly-supervised spatio-temporally grounding natural sentence in video,” in ACL , 2019
2019
Earlier work this paper cites.
Y. Wu, L. Zhu, Y. Yan, and Y. Yang, “Dual attention matching for audio-visual event localization,” in ICCV , 2019
2019
Earlier work this paper cites.
Z. Zhang, Z. Zhao, Z. Lin, J. Song, and D. Cai, “Localizing unseen activities in video via image query,” in IJCAI , 2019
2019
Earlier work this paper cites.
Y. Feng, L. Ma, W. Liu, and J. Luo, “Spatio-temporal video re-localization by warp lstm,” in CVPR , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
S. Yang, G. Li, and Y. Yu, “Dynamic graph attention for referring expression comprehension,” in ICCV , 2019
2019
Earlier work this paper cites.
J. Liang, l. Jiang, L. Cao, Y. Kalantidis, L. Li, and A. G. Hauptmann, “Focal visual-text attention for memex question answering,” IEEE TPAMI , 2019
2019
Earlier work this paper cites.
Z. Yu, D. Xu, J. Yu, T. Yu, Z. Zhao, Y. Zhuang, and D. Tao, “Activitynet-qa: A dataset for understanding complex web videos via question answering,” in AAAI , 2019
2019
Earlier work this paper cites.
H. Le, D. Sahoo, N. Chen, and S. Hoi, “Multimodal transformer networks for end-to-end video-grounded dialogue systems,” in ACL , 2019
2019
Earlier work this paper cites.
——, “Dstc7-avsd: Scene-aware video-dialogue systems with dual attention,” in AAAI workshop , 2019
2019
Earlier work this paper cites.
M. Xu, C. Zhao, D. S. Rojas, A. Thabet, and B. Ghanem, “G-tad: Sub-graph localization for temporal action detection,” in CVPR , 2020
2020
Earlier work this paper cites.
S. Zhang, H. Peng, J. Fu, and J. Luo, “Learning 2d temporal adjacent networks formoment localization with natural language,” in AAAI , vol. 34, 2020
2020
Earlier work this paper cites.
H. Zhang, A. Sun, W. Jing, and J. T. Zhou, “Span-based localizing network for natural language video localization,” in ACL , 2020
2020
Earlier work this paper cites.
J. Wu, G. Li, S. Liu, and L. Lin, “Tree-structured policy based progressive reinforcement learning for temporally language grounding in video,” in AAAI , vol. 34, 2020
2020
Earlier work this paper cites.
Z. Lin, Z. Zhao, Z. Zhang, Q. Wang, and H. Liu, “Weakly-supervised video moment retrieval via semantic completion network,” in AAAI , vol. 34, 2020
2020
Earlier work this paper cites.
Y. Yang, Z. Li, and G. Zeng, “A survey of temporal activity localization via language in untrimmed videos,” in ICCST , 2020
2020
Earlier work this paper cites.
K. Ning, M. Cai, D. Xie, and F. Wu, “An attentive sequence to sequence translator for localizing video clips by natural language,” IEEE TMM , vol. 22, 2020
2020
Earlier work this paper cites.
Y. Yuan, L. Ma, J. Wang, W. Liu, and W. Zhu, “Semantic conditioned dynamic modulation for temporal sentence grounding in videos,” IEEE TPAMI , vol. 1, 2020
2020
Earlier work this paper cites.
Z. Lin, Z. Zhao, Z. Zhang, Z. Zhang, and D. Cai, “Moment retrieval via cross-modal interaction networks with query reconstruction,” IEEE TIP , vol. 29, 2020
2020
Earlier work this paper cites.
J. Wang, L. Ma, and W. Jiang, “Temporally grounding language queries in videos by contextual boundary-aware prediction,” in AAAI , 2020
2020
Earlier work this paper cites.
X. Qu, P. Tang, Z. Zou, Y. Cheng, J. Dong, P. Zhou, and Z. Xu, “Fine-grained iterative attention network for temporal language localization in videos,” in ACM MM , 2020
2020
Earlier work this paper cites.
D. Liu, X. Qu, X.-Y. Liu, J. Dong, P. Zhou, and Z. Xu, “Jointly cross- and self-modal graph attention network for query-based moment localization,” in ACM MM , 2020
2020
Cited alongside, same era.
D. Liu, X. Qu, J. Dong, and P. Zhou, “Reasoning step-by-step: Temporal sentence localization in videos via deep rectification-modulation network,” in COLING , 2020
2020
Cited alongside, same era.
Q. Huang, J. Wei, Y. Cai, C. Zheng, J. Chen, H.-f. Leung, and Q. Li, “Aligned dual channel graph convolutional network for visual question answering,” in ACL , 2020
2020
Cited alongside, same era.
H. Wang, Z.-J. Zha, X. Chen, Z. Xiong, and J. Luo, “Dual path interaction network for video moment localization,” in ACM MM , 2020
2020
Cited alongside, same era.
J. Nam, D. Ahn, D. Kang, S. J. Ha, and J. Choi, “Zero-shot natural language video localization,” in ICCV , 2021
2021
Later among the works it cites.
J. Gao and C. Xu, “Learning video moment retrieval without a single annotated video,” IEEE TCSVT , 2021
2021
Later among the works it cites.
D. Li, R. Wu, Y. Tang, Z. Zhang, and W. Zhang, “Multi-scale 2d representation learning for weakly-supervised moment retrieval,” in ICPR , 2021
2021
Later among the works it cites.
X. Yang, F. Feng, W. Ji, M. Wang, and T.-S. Chua, “Deconfounded video moment retrieval with causal intervention,” in SIGIR , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
R. Zeng, H. Xu, W. Huang, P. Chen, M. Tan, and C. Gan, “Dense regression network for video grounding,” in CVPR , 2020
2020
Cited alongside, same era.
J. Mun, M. Cho, and B. Han, “Local-global video-text interactions for temporal grounding,” in CVPR , 2020
2020
Cited alongside, same era.
L. Chen, C. Lu, S. Tang, J. Xiao, D. Zhang, C. Tan, and X. Li, “Rethinking the bottom-up framework for query-based video localization,” in AAAI , vol. 34, 2020
2020
Cited alongside, same era.
S. Chen and Y.-G. Jiang, “Hierarchical visual-textual graph for temporal activity localization via language,” in ECCV , 2020
2020
Cited alongside, same era.
S. Chen, W. Jiang, W. Liu, and Y.-G. Jiang, “Learning modality interaction for temporal sentence localization and event captioning in videos,” in ECCV , 2020
2020
Cited alongside, same era.
C. Rodriguez, E. Marrese-Taylor, F. S. Saleh, H. LI, and S. Gould, “Proposal-free temporal moment localization of a natural-language query in video using guided attention,” in WACV , 2020
2020
Cited alongside, same era.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” IEEE TPAMI , vol. 42, 2020
2020
Cited alongside, same era.
J. Lei, T. L. Berg, and M. Bansal, “Qvhighlights: Detecting moments and highlights in videos via natural language queries,” in NeurIPS , 2021
2021
Later among the works it cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR , 2021
2021
Later among the works it cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in ICML , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
J. Lei, L. Li, L. Zhou, Z. Gan, T. L. Berg, M. Bansal, and J. Liu, “Less is more: Clipbert for video-and-language learningvia sparse sampling,” in CVPR , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
H. Xu, G. Ghosh, P.-Y. Huang, P. Arora, M. Aminzadeh, C. Feichtenhofer, F. Metze, and L. Zettlemoyer, “Vlm: Task-agnostic video-language model pre-training for video understanding,” in ACL Findings , 2021
2021
Later among the works it cites.
R. Zellers, X. Lu, J. Hessel, Y. Yu, J. S. Park, J. Cao, A. Farhadi, and Y. Choi, “Merlot: Multimodal neural script knowledge models,” in NeurIPS , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
Z. Tang, Y. Liao, S. Liu, G. Li, X. Jin, H. Jiang, Q. Yu, and D. Xu, “Human-centric spatio-temporal video grounding with visual transformers,” IEEE TCSVT , 2021
2021
Later among the works it cites.
R. Tan, B. Plummer, K. Saenko, H. Jin, and B. Russell, “Look at what i’m doing: Self-supervised spatial grounding of narrations in instructional videos,” in NeurIPS , vol. 34, 2021
2021
Later among the works it cites.
R. Su, Q. Yu, and D. Xu, “Stvgbert: A visual-linguistic transformer based framework for spatio-temporal video grounding,” in ICCV , 2021
2021
Later among the works it cites.
B. Duan, H. Tang, W. Wang, Z. Zong, G. Yang, and Y. Yan, “Audio-visual event localization via recursive fusion by joint co-attention,” in WACV , 2021
2021
Later among the works it cites.
H. Xuan, L. Luo, Z. Zhang, J. Yang, and Y. Yan, “Discriminative cross-modality attention network for temporal inconsistent audio-visual event localization,” IEEE TIP , vol. 30, 2021
2021
Later among the works it cites.
C. Xue, X. Zhong, M. Cai, H. Chen, and W. Wang, “Audio-visual event localization by learning spatial and semantic co-attention,” IEEE TMM , 2021
2021
Later among the works it cites.
L. Liu, J. Li, L. Niu, R. Xu, and L. Zhang, “Activity image-to-video retrieval by disentangling appearance and motion,” in AAAI , vol. 35, 2021
2021
Later among the works it cites.
C. Jiang, K. Huang, S. He, X. Yang, W. Zhang, X. Zhang, Y. Cheng, L. Yang, Q. Wang, F. Xu, T. Pan, and W. Chu, “Learning segment similarity and alignment in large-scale content based video retrieval,” in ACM MM , 2021
2021
Later among the works it cites.
J. Lei, T. Berg, and M. Bansal, “mTVR: Multilingual moment retrieval in videos,” in ACL , 2021
2021
Later among the works it cites.
H. Zhang, A. Sun, W. Jing, G. Nan, L. Zhen, J. T. Zhou, and R. S. M. Goh, “Video corpus moment retrieval with contrastive learning,” in SIGIR , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
S. Paul, N. C. Mithun, and A. K. Roy-Chowdhury, “Text-based localization of moments in a video corpus,” IEEE TIP , vol. 30, 2021
2021
Later among the works it cites.
Z. Hou, C.-W. Ngo, and W. K. Chan, “Conquer: Contextual query-aware ranking for video corpus moment retrieval,” in ACM MM , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
X. Sun, X. Long, D. He, S. Wen, and Z. Lian, “Vsrnet: End-to-end video segment retrieval with text query,” Pattern Recognition , vol. 119, 2021
2021
Later among the works it cites.
J. Deng, Z. Yang, T. Chen, W. Zhou, and H. Li, “Transvg: End-to-end visual grounding with transformers,” in ICCV , 2021
2021
Later among the works it cites.
X. Li, F. Zhou, C. Xu, J. Ji, and G. Yang, “Sea: Sentence encoder assembly for video retrieval by textual queries,” IEEE TMM , vol. 23, 2021
2021
Later among the works it cites.
S. Kim, S. Jeong, E. Kim, I. Kang, and N. Kwak, “Self-supervised pre-training and contrastive representation learning for multiple-choice video qa,” in AAAI , 2021
2021
Later among the works it cites.
M. Liu, L. Nie, Y. Wang, M. Wang, and Y. Rui, “A survey on video moment localization,” ACM Comput. Surv. , 2022
2022
Closest in time.
Z. Jia, M. Dong, J. Ru, L. Xue, S. Yang, and C. Li, “Stcm-net: A symmetrical one-stage network for temporal language localization in videos,” Neurocomputing , vol. 471, 2022
2022
Closest in time.
L. Zhang and R. J. Radke, “Natural language video moment localization through query-controlled temporal convolution,” in WACV , 2022
2022
Closest in time.
M. Soldan, A. Pardo, J. L. Alcázar, F. C. Heilbron, C. Zhao, S. Giancola, and B. Ghanem, “Mad: A scalable dataset for language grounding in videos from movie audio descriptions,” in CVPR , 2022
2022
Closest in time.
Y. Hu, Y. Xu, Y. Zhang, R. Feng, T. Zhang, X. Lu, and S. Gao, “Camg: Context-aware moment graph network for multimodal temporal activity localization via language,” SSRN , 2022
2022
Closest in time.
J. Gao, X. Sun, B. Ghanem, X. Zhou, and S. Ge, “Efficient video grounding with which-where reading comprehension,” IEEE TCSVT , vol. 32, 2022
2022
Closest in time.
D. Liu and W. Hu, “Skimming, locating, then perusing: A human-like framework for natural language video localization,” in ACM MM , 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
C. Guo, D. Liu, and P. Zhou, “A hybird alignment loss for temporal moment localization with natural language,” in ICME , 2022
2022
Closest in time.
B. Zhang, B. Jiang, C. Yang, and L. Pang, “Dual-channel localization networks for moment retrieval with natural language,” in ACM ICMR , 2022
2022
Closest in time.
J. Shin and J. Moon, “Learning to combine the modalities of language and video for temporal moment localization,” Computer Vision and Image Understanding , vol. 217, 2022
2022
Closest in time.
2022
Closest in time.
P. Bao and Y. Mu, “Learning sample importance for cross-scenario video temporal grounding,” in ACM ICMR , 2022
2022
Closest in time.
2022
Closest in time.
X. Ding, N. Wang, S. Zhang, Z. Huang, X. Li, M. Tang, T. Liu, and X. Gao, “Exploring language hierarchy for video grounding,” IEEE TIP , vol. 31, 2022
2022
Closest in time.
G. Wang, X. Xu, F. Shen, H. Lu, Y. Ji, and H. T. Shen, “Cross-modal dynamic networks for video moment retrieval with text query,” IEEE TMM , vol. 24, 2022
2022
Closest in time.
G. Wang, X. Jiang, N. Liu, and X. Xu, “Language-enhanced object reasoning networks for video moment retrieval with text query,” Computers and Electrical Engineering , vol. 102, 2022
2022
Closest in time.
X. Sun, X. Wang, J. Gao, Q. Liu, and X. Zhou, “You need to read again: Multi-granularity perception network for moment retrieval in videos,” in SIGIR , 2022
2022
Closest in time.
J. Li, J. Xie, L. Qian, L. Zhu, S. Tang, F. Wu, Y. Yang, Y. Zhuang, and X. Wang, “Compositional temporal grounding with structured variational cross-graph correspondence learning,” in CVPR , 2022
2022
Closest in time.
2022
Closest in time.
Z. Xu, D. Chen, K. Wei, C. Deng, and H. Xue, “Hisa: Hierarchically semantic associating for video temporal grounding,” IEEE TIP , vol. 31, 2022
2022
Closest in time.
D. Liu, X. Qu, X. Di, Y. Cheng, Z. Xu, and P. Zhou, “Memory-guided semantic learning network for temporal sentence grounding,” in AAAI , 2022
2022
Closest in time.
S. Li, C. Li, M. Zheng, and Y. Liu, “Phrase-level prediction for video temporal localization,” in ACM ICMR , 2022
2022
Closest in time.
Z. Guo, Z. Zhao, W. Jin, D. Wang, R. Liu, and J. Yu, “Taohighlight: Commodity-aware multi-modal video highlight detection in e-commerce,” IEEE TMM , vol. 24, 2022
2022
Closest in time.
D. Liu, X. Qu, P. Zhou, and Y. Liu, “Exploring motion and appearance information for temporal sentence grounding,” in AAAI , 2022
2022
Closest in time.
2022
Closest in time.
S. Yang and X. Wu, “Entity-aware and motion-aware transformers for language-driven action localization in videos,” in IJCAI , 2022
2022
Closest in time.
X. Shen, L. Lan, H. Tan, X. Zhang, X. Ma, and Z. Luo, “Joint modality synergy and spatio-temporal cue purification for moment localization,” in ACM ICMR , 2022
2022
Closest in time.
H. Fu and H. Wang, “Multiple cross-attention for video-subtitle moment retrieval,” Pattern Recognition Letters , vol. 156, 2022
2022
Closest in time.
L. Zhang and R. J. Radke, “Natural language video moment localization through query-controlled temporal convolution,” in WACV , 2022
2022
Closest in time.
Y. Zeng, “Point prompt tuning for temporally language grounding,” in SIGIR , 2022
2022
Closest in time.
J. Hao, H. Sun, P. Ren, J. Wang, Q. Qi, and J. Liao, “Query-aware video encoder for video moment retrieval,” Neurocomputing , vol. 483, 2022
2022
Closest in time.
D. Liu, X. Qu, and W. Hu, “Reducing the vision and language bias for temporal sentence grounding,” in ACM MM , 2022
2022
Closest in time.
Y. Xu, Y. Zhang, R. Feng, R.-W. Zhao, T. Zhang, X. Lu, and S. Gao, “Stdnet: Spatio-temporal decomposed network for video grounding,” in ICME , 2022
2022
Closest in time.
B. Li, Y. Weng, B. Sun, and S. Li, “Towards visual-prompt temporal answering grounding in medical instructional video,” in ACM MM , 2022
2022
Closest in time.
2022
Closest in time.
Y. Zeng, D. Cao, S. Lu, H. Zhang, J. Xu, and Z. Qin, “Moment is important: Language-based video moment retrieval via adversarial learning,” ACM TMCCA , vol. 18, 2022
2022
Closest in time.
H. Jiang and Y. Mu, “Joint video summarization and moment localization by cross-task sample transfer,” in CVPR , 2022
2022
Closest in time.
X. Jiang, X. Xu, J. Zhang, F. Shen, Z. Cao, and X. Cai, “Gtlr: Graph-based transformer with language reconstruction for video paragraph grounding,” in ICME , 2022
2022
Closest in time.
X. Jiang, X. Xu, J. Zhang, F. Shen, Z. Cao, and H. T. Shen, “Semi-supervised video paragraph grounding with contrastive encoder,” in CVPR , 2022
2022
Closest in time.
X. Yang, S. Wang, J. Dong, J. Dong, M. Wang, and T.-S. Chua, “Video moment retrieval with cross-modal neural architecture search,” IEEE TIP , vol. 31, 2022
2022
Closest in time.
M. Cao, T. Yang, J. Weng, C. Zhang, J. Wang, and Y. Zou, “Locvtp: Video-text pre-training for temporal localization,” in ECCV , 2022
2022
Closest in time.
H. S. Nawaz, Z. Shi, Y. Gan, A. Hirpa, J. Dong, and H. Zheng, “Temporal moment localization via natural language by utilizing video question answers as a special variant and bypassing nlp for corpora,” IEEE TCSVT , vol. 32, 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
Y. Liu, S. Li, Y. Wu, C. W. Chen, Y. Shan, and X. Qie, “Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection,” in CVPR , 2022
2022
Closest in time.
J. Chen, W. Luo, W. Zhang, and L. Ma, “Explore inter-contrast between videos via composition for weakly supervised temporal sentence grounding,” in AAAI , 2022
2022
Closest in time.
2022
Closest in time.
Y. Wang, M. Liu, Y. Wei, Z. Cheng, Y. Wang, and L. Nie, “Siamese alignment network for weakly supervised video moment retrieval,” IEEE TMM , 2022
2022
Closest in time.
T. Han, K. Wang, J. Yu, and J. Fan, “Weakly supervised moment localization with natural language based on semantic reconstruction,” Image and Vision Computing , vol. 126, 2022
2022
Closest in time.
M. Zheng, Y. Huang, Q. Chen, Y. Peng, and Y. Liu, “Weakly supervised temporal sentence grounding with gaussian-based contrastive proposal learning,” in CVPR , 2022
2022
Closest in time.
M. Zheng, Y. Huang, Q. Chen, and Y. Liu, “Weakly supervised video moment localization with contrastive negative sample mining,” in AAAI , 2022
2022
Closest in time.
D. Liu, X. Qu, Y. Wang, X. Di, K. Zou, Y. Cheng, Z. Xu, and P. Zhou, “Unsupervised temporal video grounding with deep semantic clustering,” in AAAI , 2022
2022
Closest in time.
S. Paul, N. C. Mithun, and A. K. Roy-Chowdhury, “Text-based temporal localization of novel events,” ArXiv , 2022
2022
Closest in time.
Z. Xu, K. Wei, X. Yang, and C. Deng, “Point-supervised video temporal grounding,” IEEE TMM , 2022
2022
Closest in time.
R. Cui, T. Qian, P. Peng, E. Daskalaki, J. Chen, X.-W. Guo, H. Sun, and Y.-G. Jiang, “Video moment retrieval from text queries via single frame annotation,” in SIGIR , 2022
2022
Closest in time.
H. Zhou, C. Zhang, Y. Luo, C. Hu, and W. Zhang, “Thinking inside uncertainty: Interest moment perception for diverse temporal grounding,” IEEE TCSVT , vol. 32, 2022
2022
Closest in time.
M. Cao, J. Jiang, L. Chen, and Y. Zou, “Correspondence matters for video referring expression comprehension,” in ACM MM , 2022
2022
Closest in time.
Y. Li, J. Yu, Z. Cai, and Y. Pan, “Cross-modal target retrieval for tracking by natural language,” in CVPR Workshops , 2022
2022
Closest in time.
M. Li, T. Wang, H. Zhang, S. Zhang, Z. Zhao, J. Miao, W. Zhang, W. Tan, J. Wang, P. Wang, S. Pu, and F. Wu, “End-to-end modeling via information tree for one-shot natural language spatial video grounding,” in ACL , 2022
2022
Closest in time.
2022
Closest in time.
Z. Lin, C. Tan, J. Hu, Z. Jin, T. Ye, and W. Zheng, “Stvgformer: Spatio-temporal video grounding with static-dynamic cross-modal understanding,” in ACM MM Workshop , 2022
2022
Closest in time.
A. Yang, A. Miech, J. Sivic, I. Laptev, and C. Schmid, “Tubedetr: Spatio-temporal video grounding with transformers,” in CVPR , 2022
2022
Closest in time.
Y. Xia, Z. Zhao, S. Ye, Y. Zhao, H. Li, and Y. Ren, “Video-guided curriculum learning for spoken video grounding,” in ACM MM , 2022
2022
Closest in time.
S. Yoon, D. Kim, J. Kim, and C. D. Yoo, “Cascaded mpn: Cascaded moment proposal network for video corpus moment retrieval,” IEEE Access , vol. 10, 2022
2022
Closest in time.
J. Liu, T. Yu, H. Peng, M. Sun, and P. Li, “Cross-lingual cross-modal consolidation for effective multilingual video corpus moment retrieval,” in Findings of NAACL , 2022
2022
Closest in time.
D. Kim, S. Yoon, J. W. Hong, and C. D. Yoo, “Semantic association network for video corpus moment retrieval,” in ICASSP , 2022
2022
Closest in time.
Z. Wang, Y. Wu, K. Narasimhan, and O. Russakovsky, “Multi-query video retrieval,” in ECCV , 2022
2022
Closest in time.
X. Wang, L. Zhu, Z. Zheng, M. Xu, and Y. Yang, “Align and tell: Boosting text-video retrieval with local alignment and fine-grained supervision,” IEEE TMM , 2022
2022
Closest in time.