Fetching the paper…
Reading the bibliography…
Temporal sentence grounding aims to localize a target segment in an untrimmed video semantically according to a given sentence query.
M. Regneri, M. Rohrbach, D. Wetzel, S. Thater, B. Schiele, and M. Pinkal, “Grounding action descriptions in videos,” TACL , vol. 1, pp. 25–36, 2013
2013
Earlier work this paper cites.
R. Socher, A. Karpathy, Q. V. Le, C. D. Manning, and A. Y. Ng, “Grounded compositional semantics for finding and describing images with sentences,” TACL , vol. 2, pp. 207–218, 2014
2014
Earlier work this paper cites.
D. Lin, S. Fidler, C. Kong, and R. Urtasun, “Visual semantic search: Retrieving videos via complex textual queries,” in CVPR , 2014, pp. 2657–2664
2014
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” in NIPS , 2014, pp. 568–576
2014
Earlier work this paper cites.
J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation,” in EMNLP , 2014, pp. 1532–1543
2014
Earlier work this paper cites.
J. Chung, C. Gulcehre, K. Cho, and Y. Bengio, “Empirical evaluation of gated recurrent neural networks on sequence modeling,” in NIPS , 2014
2014
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” in NIPS , 2015, pp. 91–99
2015
Earlier work this paper cites.
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learning spatiotemporal features with 3d convolutional networks,” in ICCV , 2015, pp. 4489–4497
2015
Earlier work this paper cites.
R. Xu, C. Xiong, W. Chen, and J. Corso, “Jointly modeling deep video and compositional text to bridge vision and language in a unified framework,” in AAAI , vol. 29, no. 1, 2015
2015
Earlier work this paper cites.
F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. Carlos Niebles, “Activitynet: A large-scale video benchmark for human activity understanding,” in CVPR , 2015, pp. 961–970
2015
Earlier work this paper cites.
Z. Shou, D. Wang, and S.-F. Chang, “Temporal action localization in untrimmed videos via multi-stage cnns,” in CVPR , 2016, pp. 1049–1058
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016, pp. 770–778
2016
Earlier work this paper cites.
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy, “Generation and comprehension of unambiguous object descriptions,” in CVPR , 2016, pp. 11–20
2016
Earlier work this paper cites.
R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell, “Natural language object retrieval,” in CVPR , 2016, pp. 4555–4564
2016
Earlier work this paper cites.
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele, “Grounding of textual phrases in images by reconstruction,” in ECCV , 2016, pp. 817–834
2016
Earlier work this paper cites.
M. Otani, Y. Nakashima, E. Rahtu, J. Heikkilä, and N. Yokoya, “Learning joint representations of videos and sentences with web image search,” in ECCV , 2016, pp. 651–667
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta, “Hollywood in homes: Crowdsourcing data collection for activity understanding,” in ECCV , 2016, pp. 510–526
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
J. Lu, J. Yang, D. Batra, and D. Parikh, “Hierarchical question-image co-attention for visual question answering,” in NIPS , 2016, pp. 289–297
2016
Earlier work this paper cites.
Y. Zhao, Y. Xiong, L. Wang, Z. Wu, X. Tang, and D. Lin, “Temporal action detection with structured segment networks,” in ICCV , 2017, pp. 2914–2923
2017
Earlier work this paper cites.
L. Anne Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell, “Localizing moments in video with natural language,” in ICCV , 2017
2017
Earlier work this paper cites.
J. Gao, C. Sun, Z. Yang, and R. Nevatia, “Tall: Temporal activity localization via language query,” in ICCV , 2017, pp. 5267–5275
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in ICCV , 2017, pp. 2961–2969
2017
Earlier work this paper cites.
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. Carlos Niebles, “Dense-captioning events in videos,” in ICCV , 2017, pp. 706–715
2017
Earlier work this paper cites.
T. Lin, X. Zhao, and Z. Shou, “Single shot temporal action detection,” in ACM MM , 2017, pp. 988–996
2017
Earlier work this paper cites.
H. Xu, A. Das, and K. Saenko, “R-c3d: Region convolutional 3d network for temporal activity detection,” in ICCV , 2017, pp. 5783–5792
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NIPS , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in CVPR , 2017, pp. 6299–6308
2017
Earlier work this paper cites.
W. Jiang, L. Ma, Y.-G. Jiang, W. Liu, and T. Zhang, “Recurrent fusion network for image captioning,” in ECCV , 2018, pp. 499–515
2018
Earlier work this paper cites.
J. Chen, X. Chen, L. Ma, Z. Jie, and T.-S. Chua, “Temporally grounding natural sentence in video,” in EMNLP , 2018, pp. 162–171
2018
Earlier work this paper cites.
Y.-W. Chao, S. Vijayanarasimhan, B. Seybold, D. A. Ross, J. Deng, and R. Sukthankar, “Rethinking the faster r-cnn architecture for temporal action localization,” in CVPR , 2018, pp. 1130–1139
2018
Cited alongside, same era.
T. Lin, X. Zhao, H. Su, C. Wang, and M. Yang, “Bsn: Boundary sensitive network for temporal action proposal generation,” in ECCV , 2018, pp. 3–19
2018
Cited alongside, same era.
M. Liu, X. Wang, L. Nie, X. He, B. Chen, and T.-S. Chua, “Attentive moment retrieval in videos,” in SIGIR , 2018, pp. 15–24
2018
Cited alongside, same era.
Q. Li, Z. Han, and X.-M. Wu, “Deeper insights into graph convolutional networks for semi-supervised learning,” in AAAI , 2018
2018
Cited alongside, same era.
L. Gao, P. Zeng, J. Song, Y.-F. Li, W. Liu, T. Mei, and H. T. Shen, “Structured two-stream attention network for video question answering,” in AAAI , vol. 33, no. 01, 2019, pp. 6391–6398
L. Yang, H. Peng, D. Zhang, J. Fu, and J. Han, “Revisiting anchor mechanisms for temporal action localization,” IEEE TIP , vol. 29, pp. 8535–8548, 2020
2020
Later among the works it cites.
J. Wang, L. Ma, and W. Jiang, “Temporally grounding language queries in videos by contextual boundary-aware prediction,” in AAAI , 2020
2020
Later among the works it cites.
J. Wu, G. Li, S. Liu, and L. Lin, “Tree-structured policy based progressive reinforcement learning for temporally language grounding in video,” in AAAI , vol. 34, no. 07, 2020, pp. 12 386–12 393
2020
Later among the works it cites.
R. Zeng, H. Xu, W. Huang, P. Chen, M. Tan, and C. Gan, “Dense regression network for video grounding,” in CVPR , 2020, pp. 10 287–10 296
2020
Later among the works it cites.
X. Qu, P. Tang, Z. Zou, Y. Cheng, J. Dong, P. Zhou, and Z. Xu, “Fine-grained iterative attention network for temporal language localization in videos,” in ACM MM , 2020, pp. 4280–4288
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
Z. Zhang, Z. Lin, Z. Zhao, and Z. Xiao, “Cross-modal interaction networks for query-based moment retrieval in videos,” in SIGIR , 2019, pp. 655–664
2019
Cited alongside, same era.
Y. Yuan, T. Mei, and W. Zhu, “To find where you talk: Temporal sentence localization in video with attention based location regression,” in AAAI , vol. 33, 2019, pp. 9159–9166
2019
Cited alongside, same era.
2019
Cited alongside, same era.
F. Long, T. Yao, Z. Qiu, X. Tian, J. Luo, and T. Mei, “Gaussian temporal awareness networks for action localization,” in CVPR , 2019, pp. 344–353
2019
Cited alongside, same era.
T. Lin, X. Liu, X. Li, E. Ding, and S. Wen, “Bmn: Boundary-matching network for temporal action proposal generation,” in ICCV , 2019, pp. 3889–3898
2019
Cited alongside, same era.
H. Xu, A. Das, and K. Saenko, “Two-stream region convolutional 3d network for temporal activity detection,” IEEE TPAMI , vol. 41, no. 10, pp. 2319–2332, 2019
2019
Cited alongside, same era.
H. Xu, K. He, B. A. Plummer, L. Sigal, S. Sclaroff, and K. Saenko, “Multilevel language and vision integration for text-to-clip retrieval,” in AAAI , vol. 33, 2019, pp. 9062–9069
2019
Cited alongside, same era.
2020
Later among the works it cites.
Y. Wang, J. Deng, W. Zhou, and H. Li, “Weakly supervised temporal adjacent network for language grounding,” IEEE TMM , 2021
2021
Later among the works it cites.
H. Tang, J. Zhu, M. Liu, Z. Gao, and Z. Cheng, “Frame-wise cross-modal matching for video moment retrieval,” IEEE TMM , 2021
2021
Later among the works it cites.
C. Sun, H. Song, X. Wu, Y. Jia, and J. Luo, “Exploiting informative video segments for temporal action localization,” IEEE TMM , 2021
2021
Later among the works it cites.
J. Wang, B. Bao, and C. Xu, “Dualvgr: A dual-visual graph reasoning unit for video question answering,” IEEE TMM , 2021
2021
Later among the works it cites.
D. Liu, X. Qu, J. Dong, and P. Zhou, “Adaptive proposal generation network for temporal sentence localization in videos,” in EMNLP , 2021, pp. 9292–9301
2021
Later among the works it cites.
D. Liu, X. Qu, and P. Zhou, “Progressively guide to attend: An iterative alignment framework for temporal sentence grounding,” in EMNLP , 2021, pp. 9302–9311
2021
Later among the works it cites.
Y. Zeng, D. Cao, X. Wei, M. Liu, Z. Zhao, and Z. Qin, “Multi-modal relational graph for cross-modal video moment retrieval,” in CVPR , 2021, pp. 2215–2224
2021
Later among the works it cites.
J. Kangaspunta, A. Piergiovanni, R. Jonschkowski, M. Ryoo, and A. Angelova, “Adaptive intermediate representations for video understanding,” in CVPR , 2021, pp. 1602–1612
2021
Later among the works it cites.
A. Blattmann, T. Milbich, M. Dorkenwald, and B. Ommer, “Understanding object dynamics for interactive image-to-video synthesis,” in CVPR , 2021, pp. 5171–5181
2021
Later among the works it cites.
M. Dorkenwald, T. Milbich, A. Blattmann, R. Rombach, K. G. Derpanis, and B. Ommer, “Stochastic image-to-video synthesis using cinns,” in CVPR , 2021, pp. 3742–3753
2021
Later among the works it cites.
S. Woo, D. Kim, J.-Y. Lee, and I. S. Kweon, “Learning to associate every segment for video panoptic segmentation,” in CVPR , 2021, pp. 2705–2714
2021
Later among the works it cites.
V. Jayasundara, D. Roy, and B. Fernando, “Flowcaps: Optical flow estimation with capsule networks for action recognition,” in WACV , 2021, pp. 3409–3418
2021
Later among the works it cites.
C. Bai, H. Li, J. Zhang, L. Huang, and L. Zhang, “Unsupervised adversarial instance-level image retrieval,” IEEE TMM , 2021
2021
Later among the works it cites.
W. Chen, Y. Liu, N. Pu, W. Wang, L. Liu, and M. S. Lew, “Feature estimations based correlation distillation for incremental image retrieval,” IEEE TMM , 2021
2021
Later among the works it cites.
Z. Weng and Y. Zhu, “Online hashing with bit selection for image retrieval,” IEEE TMM , vol. 23, pp. 1868–1881, 2021
2021
Later among the works it cites.
X. Song, J. Chen, Z. Wu, and Y.-G. Jiang, “Spatial-temporal graphs for cross-modal text2video retrieval,” IEEE TMM , 2021
2021
Later among the works it cites.
M. Qi, J. Qin, Y. Yang, Y. Wang, and J. Luo, “Semantics-aware spatial-temporal binaries for cross-modal video retrieval,” IEEE TIP , vol. 30, pp. 2989–3004, 2021
2021
Later among the works it cites.
R. R. A. Pramono, Y.-T. Chen, and W.-H. Fang, “Spatial-temporal action localization with hierarchical self-attention,” IEEE TMM , 2021
2021
Later among the works it cites.
Y. Zhai, L. Wang, W. Tang, Q. Zhang, N. Zheng, and G. Hua, “Action coherence network for weakly-supervised temporal action localization,” IEEE TMM , 2021
2021
Later among the works it cites.
D. Liu, X. Qu, J. Dong, P. Zhou, Y. Cheng, W. Wei, Z. Xu, and Y. Xie, “Context-aware biaffine localizing network for temporal sentence grounding,” in CVPR , 2021
2021
Later among the works it cites.
S. Xiao, L. Chen, S. Zhang, W. Ji, J. Shao, L. Ye, and J. Xiao, “Boundary proposal network for two-stage natural language video localization,” in AAAI , 2021
2021
Later among the works it cites.
G. Nan, R. Qiao, Y. Xiao, J. Liu, S. Leng, H. Zhang, and W. Lu, “Interventional video grounding with dual contrastive learning,” in CVPR , 2021
2021
Later among the works it cites.
D. Liu, X. Qu, P. Zhou, and Y. Liu, “Exploring motion and appearance information for temporal sentence grounding,” in AAAI , 2022
2022
Closest in time.
D. Liu, X. Qu, X. Di, Y. Cheng, Z. X. Xu, and P. Zhou, “Memory-guided semantic learning network for temporal sentence grounding,” in AAAI , 2022
2022
Closest in time.
D. Liu, X. Qu, Y. Wang, X. Di, K. Zou, Y. Cheng, Z. Xu, and P. Zhou, “Unsupervised temporal video grounding with deep semantic clustering,” in AAAI , 2022
2022
Closest in time.
Z. Liu, S. Wu, S. Jin, Q. Liu, S. Ji, S. Lu, and L. Cheng, “Investigating pose representations and motion contexts modeling for 3d motion prediction,” IEEE TPAMI , 2022
2022
Closest in time.