Fetching the paper…
Reading the bibliography…
We address the challenging task of cross-modal moment retrieval, which aims to localize a temporal segment from an untrimmed video described by a natural language query.
Heterogeneous memory enhanced multimodal attention model for video question answering
Fan, C.; Zhang, X.; Zhang, S.; Wang, W.; Zhang, C.; and Huang, H. 2019 · 2007
Earlier work this paper cites.
Script data for attribute-based recognition of composite activities
Rohrbach, M.; Regneri, M.; Andriluka, M.; Amin, S.; Pinkal, M.; and Schiele, B. 2012 · 2012
Earlier work this paper cites.
Grounding action descriptions in videos
Regneri, M.; Rohrbach, M.; Wetzel, D.; Thater, S.; Schiele, B.; and Pinkal, M. 2013 · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J.; Socher, R.; and Manning, C. D. 2014 · 2014
Earlier work this paper cites.
Skip-thought vectors
Kiros, R.; Zhu, Y.; Salakhutdinov, R. R.; Zemel, R.; Urtasun, R.; Torralba, A.; and Fidler, S. 2015 · 2015
Earlier work this paper cites.
Very Deep Convolutional Networks for Large-Scale Image Recognition
Simonyan, K.; and Zisserman, A. 2015 · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Tran, D.; Bourdev, L.; Fergus, R.; Torresani, L.; and Paluri, M. 2015 · 2015
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Sigurdsson, G. A.; Varol, G.; Wang, X.; Farhadi, A.; Laptev, I.; and Gupta, A. 2016 · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
Anne Hendricks, L.; Wang, O.; Shechtman, E.; Sivic, J.; Darrell, T.; and Russell, B. 2017 · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J.; and Zisserman, A. 2017 · 2017
Earlier work this paper cites.
Tall: Temporal activity localization via language query
Gao, J.; Sun, C.; Yang, Z.; and Nevatia, R. 2017 · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Paszke, A.; Gross, S.; Chintala, S.; Chanan, G.; Yang, E.; DeVito, Z.; Lin, Z.; Desmaison, A.; Antiga, L.; and Lerer, A. 2017 · 2017
Cited alongside, same era.
Temporally grounding natural sentence in video
Chen, J.; Chen, X.; Ma, L.; Jie, Z.; and Chua, T.-S. 2018 · 2018
Cited alongside, same era.
Film: Visual reasoning with a general conditioning layer
Perez, E.; Strub, F.; De Vries, H.; Dumoulin, V.; and Courville, A. 2018 · 2018
Cited alongside, same era.
Val: Visual-attention action localizer
Song, X.; and Han, Y. 2018 · 2018
Cited alongside, same era.
Multi-modal Circulant Fusion for Video-to-Language and Backward
Wu, A.; and Han, Y. 2018 · 2018
Cited alongside, same era.
Cross-modal video moment retrieval with spatial and language-temporal attention
Jiang, B.; Huang, X.; Yang, C.; and Yuan, J. 2019 · 2019
Later among the works it cites.
Language-driven temporal activity localization: A semantic matching reinforcement learning model
Wang, W.; Huang, Y.; and Wang, L. 2019 · 2019
Later among the works it cites.
YouMakeup: A Large-Scale Domain-Specific Multimodal Dataset for Fine-Grained Semantic Comprehension
Wang, W.; Wang, Y.; Chen, S.; and Jin, Q. 2019 · 2019
Later among the works it cites.
Multilevel language and vision integration for text-to-clip retrieval
Xu, H.; He, K.; Plummer, B. A.; Sigal, L.; Sclaroff, S.; and Saenko, K. 2019 · 2019
Later among the works it cites.
Semantic Conditioned Dynamic Modulation for Temporal Sentence Grounding in Videos
Yuan, Y.; Ma, L.; Wang, J.; Liu, W.; and Zhu, W. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Semantic proposal for activity localization in videos via sentence query
Chen, S.; and Jiang, Y.-G. 2019 · 2019
Cited alongside, same era.
Mac: Mining activity concepts for language-based temporal localization
Ge, R.; Gao, J.; Chen, K.; and Nevatia, R. 2019 · 2019
Cited alongside, same era.
Tripping through time: Efficient localization of activities in videos
Hahn, M.; Kadav, A.; Rehg, J. M.; and Graf, H. P. 2019 · 2019
Cited alongside, same era.
Read, watch, and move: Reinforcement learning for temporally grounding natural language descriptions in videos
He, D.; Zhao, X.; Huang, J.; Li, F.; Liu, X.; and Wen, S. 2019 · 2019
Cited alongside, same era.
Attentive moment retrieval in videos
Liu, M.; Wang, X.; Nie, L.; He, X.; Chen, B.; and Chua, T.-S. 2018a
Cited in the paper.
Cross-modal moment localization in videos
Liu, M.; Wang, X.; Nie, L.; Tian, Q.; Chen, B.; and Chua, T.-S. 2018b
Cited in the paper.
To find where you talk: Temporal sentence localization in video with attention based location regression
Yuan, Y.; Mei, T.; and Zhu, W. 2019 · 2019
Later among the works it cites.
Tree-Structured Policy based Progressive Reinforcement Learning for Temporally Language Grounding in Video
Wu, J.; Li, G.; Liu, S.; and Lin, L. 2020 · 2020
Closest in time.
Dense regression network for video grounding
Zeng, R.; Xu, H.; Huang, W.; Chen, P.; Tan, M.; and Gan, C. 2020 · 2020
Closest in time.
Learning 2D Temporal Adjacent Networks for Moment Localization with Natural Language
Zhang, S.; Peng, H.; Fu, J.; and Luo, J. 2020 · 2020
Closest in time.