Fetching the paper…
Reading the bibliography…
Video grounding aims to locate a moment of interest matching the given query sentence from an untrimmed video.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M.; and Hyvärinen, A. 2010 · 2010
Earlier work this paper cites.
Grounding Action Descriptions in Videos
Regneri, M.; Rohrbach, M.; Wetzel, D.; Thater, S.; Schiele, B.; and Pinkal, M. 2013 · 2013
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Caba Heilbron, F.; Escorcia, V.; Ghanem, B.; and Carlos Niebles, J. 2015 · 2015
Earlier work this paper cites.
Localizing moments in video with natural language
Anne Hendricks, L.; Wang, O.; Shechtman, E.; Sivic, J.; Darrell, T.; and Russell, B. 2017 · 2017
Earlier work this paper cites.
Soft-NMS–improving object detection with one line of code
Bodla, N.; Singh, B.; Chellappa, R.; and Davis, L. S. 2017 · 2017
Earlier work this paper cites.
Tall: Temporal activity localization via language query
Gao, J.; Sun, C.; Yang, Z.; and Nevatia, R. 2017 · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Krishna, R.; Hata, K.; Ren, F.; Fei-Fei, L.; and Carlos Niebles, J. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Temporally grounding natural sentence in video
Chen, J.; Chen, X.; Ma, L.; Jie, Z.; and Chua, T.-S. 2018 · 2018
Earlier work this paper cites.
Jointly localizing and describing events for dense video captioning
Li, Y.; Yao, T.; Pan, Y.; Chao, H.; and Mei, T. 2018 · 2018
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2018 · 2018
Earlier work this paper cites.
Bidirectional attentive fusion with context gating for dense video captioning
Wang, J.; Jiang, W.; Ma, L.; Liu, W.; and Xu, Y. 2018 · 2018
Earlier work this paper cites.
Semantic proposal for activity localization in videos via sentence query
Chen, S.; and Jiang, Y.-G. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Kenton, J. D. M.-W. C.; and Toutanova, L. K. 2019 · 2019
Cited alongside, same era.
Debug: A dense bottom-up grounding approach for natural language video localization
Lu, C.; Chen, L.; Tan, C.; Li, X.; and Xiao, J. 2019 · 2019
Cited alongside, same era.
To find where you talk: Temporal sentence localization in video with attention based location regression
Yuan, Y.; Mei, T.; and Zhu, W. 2019 · 2019
Cited alongside, same era.
Rethinking the bottom-up framework for query-based video localization
Chen, L.; Lu, C.; Tang, S.; Xiao, J.; Zhang, D.; Tan, C.; and Li, X. 2020 · 2020
Cited alongside, same era.
Supervised contrastive learning
Interaction-integrated network for natural language moment localization
Ning, K.; Xie, L.; Liu, J.; Wu, F.; and Tian, Q. 2021 · 2021
Later among the works it cites.
End-to-end dense video captioning with parallel decoding
Wang, T.; Zhang, R.; Lu, Z.; Zheng, F.; Cheng, R.; and Luo, P. 2021 · 2021
Later among the works it cites.
Boundary proposal network for two-stage natural language video localization
Xiao, S.; Chen, L.; Zhang, S.; Ji, W.; Shao, J.; Ye, L.; and Xiao, J. 2021 · 2021
Later among the works it cites.
Cascaded prediction network via segment tree for temporal video grounding
Zhao, Y.; Zhao, Z.; Zhang, Z.; and Lin, Z. 2021 · 2021
Later among the works it cites.
Compositional temporal grounding with structured variational cross-graph correspondence learning
Li, J.; Xie, J.; Qian, L.; Zhu, L.; Tang, S.; Wu, F.; Yang, Y.; Zhuang, Y.; and Wang, X. E. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Khosla, P.; Teterwak, P.; Wang, C.; Sarna, A.; Tian, Y.; Isola, P.; Maschinot, A.; Liu, C.; and Krishnan, D. 2020 · 2020
Cited alongside, same era.
Jointly cross-and self-modal graph attention network for query-based moment localization
Liu, D.; Qu, X.; Liu, X.-Y.; Dong, J.; Zhou, P.; and Xu, Z. 2020 · 2020
Cited alongside, same era.
Local-global video-text interactions for temporal grounding
Mun, J.; Cho, M.; and Han, B. 2020 · 2020
Cited alongside, same era.
An efficient framework for dense video captioning
Suin, M.; and Rajagopalan, A. 2020 · 2020
Cited alongside, same era.
Dense regression network for video grounding
Zeng, R.; Xu, H.; Huang, W.; Chen, P.; Tan, M.; and Gan, C. 2020 · 2020
Cited alongside, same era.
Learning 2d temporal adjacent networks for moment localization with natural language
Zhang, S.; Peng, H.; Fu, J.; and Luo, J. 2020 · 2020
Cited alongside, same era.
Correspondence matters for video referring expression comprehension
Cao, M.; Jiang, J.; Chen, L.; and Zou, Y. 2022a
Cited in the paper.
Liu, D.; and Hu, W. 2022 · 2022
Later among the works it cites.
Unsupervised pre-training for temporal action localization tasks
Zhang, C.; Yang, T.; Weng, J.; Cao, M.; Wang, J.; and Zou, Y. 2022 · 2022
Later among the works it cites.
Iterative Proposal Refinement for Weakly-Supervised Video Grounding
Cao, M.; Wei, F.; Xu, C.; Geng, X.; Chen, L.; Zhang, C.; Zou, Y.; Shen, T.; and Jiang, D. 2023 · 2023
Closest in time.
G2l: Semantically aligned and uniform video grounding via geodesic and game theory
Li, H.; Cao, M.; Cheng, X.; Li, Y.; Zhu, Z.; and Zou, Y. 2023 · 2023
Closest in time.
Improving Reference-based Distinctive Image Captioning with Contrastive Rewards
Mao, Y.; Xiao, J.; Zhang, D.; Cao, M.; Shao, J.; Zhuang, Y.; and Chen, L. 2023 · 2023
Closest in time.
Concept-Aware Video Captioning: Describing Videos With Effective Prior Information
Yang, B.; Cao, M.; and Zou, Y. 2023 · 2023
Closest in time.