Fetching the paper…
Reading the bibliography…
The task of language-guided video temporal grounding is to localize the particular video clip corresponding to a query sentence in an untrimmed video.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z.; Yang, Z.; Yang, Y.; Carbonell, J.; Le, Q. V.; and Salakhutdinov, R. 2019 · 1901
Earlier work this paper cites.
Excl: Extractive clip localization using natural language descriptions
Ghosh, S.; Agarwal, A.; Parekh, Z.; and Hauptmann, A. 2019 · 1904
Earlier work this paper cites.
Long short-term memory
Hochreiter, S.; and Schmidhuber, J. 1997 · 1997
Earlier work this paper cites.
A Better Use of Audio-Visual Cues: Dense Video Captioning with Bi-modal Transformer
Iashin, V.; and Rahtu, E. 2020 · 2005
Earlier work this paper cites.
Efficient non-maximum suppression
Neubeck, A.; and Van Gool, L. 2006 · 2006
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J.; Socher, R.; and Manning, C. D. 2014 · 2014
Earlier work this paper cites.
Fast r-cnn
Girshick, R. 2015 · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Tran, D.; Bourdev, L.; Fergus, R.; Torresani, L.; and Paluri, M. 2015 · 2015
Earlier work this paper cites.
Localizing moments in video with natural language
Anne Hendricks, L.; Wang, O.; Shechtman, E.; Sivic, J.; Darrell, T.; and Russell, B. 2017 · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J.; and Zisserman, A. 2017 · 2017
Earlier work this paper cites.
Coarse-to-fine question answering for long documents
Choi, E.; Hewlett, D.; Uszkoreit, J.; Polosukhin, I.; Lacoste, A.; and Berant, J. 2017 · 2017
Earlier work this paper cites.
Tall: Temporal activity localization via language query
Gao, J.; Sun, C.; Yang, Z.; and Nevatia, R. 2017 · 2017
Cited alongside, same era.
Dense-captioning events in videos
Krishna, R.; Hata, K.; Ren, F.; Fei-Fei, L.; and Carlos Niebles, J. 2017 · 2017
Cited alongside, same era.
A structured self-attentive sentence embedding
Lin, Z.; Feng, M.; Santos, C. N. d.; Yu, M.; Xiang, B.; Zhou, B.; and Bengio, Y. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Attentive moment retrieval in videos
Liu, M.; Wang, X.; Nie, L.; He, X.; Chen, B.; and Chua, T.-S. 2018 · 2018
Cited alongside, same era.
Learning joint embedding with multimodal cues for cross-modal video-text retrieval
TS-LSTM and temporal-inception: Exploiting spatiotemporal dynamics for activity recognition
Ma, C.-Y.; Chen, M.-H.; Kira, Z.; and AlRegib, G. 2019 · 2019
Later among the works it cites.
Fcos: Fully convolutional one-stage object detection
Tian, Z.; Shen, C.; Chen, H.; and He, T. 2019 · 2019
Later among the works it cites.
Language-driven temporal activity localization: A semantic matching reinforcement learning model
Wang, W.; Huang, Y.; and Wang, L. 2019 · 2019
Later among the works it cites.
Semantic conditioned dynamic modulation for temporal sentence grounding in videos
Yuan, Y.; Ma, L.; Wang, J.; Liu, W.; and Zhu, W. 2019 · 2019
Later among the works it cites.
To find where you talk: Temporal sentence localization in video with attention based location regression
Yuan, Y.; Mei, T.; and Zhu, W. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mithun, N. C.; Li, J.; Metze, F.; and Roy-Chowdhury, A. K. 2018 · 2018
Cited alongside, same era.
Eco: Efficient convolutional network for online video understanding
Zolfaghari, M.; Singh, K.; and Brox, T. 2018 · 2018
Cited alongside, same era.
Murel: Multimodal relational reasoning for visual question answering
Cadene, R.; Ben-Younes, H.; Cord, M.; and Thome, N. 2019 · 2019
Cited alongside, same era.
Multi-modality latent interaction network for visual question answering
Gao, P.; You, H.; Zhang, Z.; Wang, X.; and Li, H. 2019 · 2019
Cited alongside, same era.
Video action transformer network
Girdhar, R.; Carreira, J.; Doersch, C.; and Zisserman, A. 2019 · 2019
Cited alongside, same era.
Read, watch, and move: Reinforcement learning for temporally grounding natural language descriptions in videos
He, D.; Zhao, X.; Huang, J.; Li, F.; Liu, X.; and Wen, S. 2019 · 2019
Cited alongside, same era.
End-to-end dense video captioning with masked transformer
Zhou, L.; Zhou, Y.; Corso, J. J.; Socher, R.; and Xiong, C. 2018a
Cited in the paper.
Chen, S.; Zhao, Y.; Jin, Q.; and Wu, Q. 2020 · 2020
Closest in time.
Local-Global Video-Text Interactions for Temporal Grounding
Mun, J.; Cho, M.; and Han, B. 2020 · 2020
Closest in time.
Proposal-free temporal moment localization of a natural-language query in video using guided attention
Rodriguez, C.; Marrese-Taylor, E.; Saleh, F. S.; Li, H.; and Gould, S. 2020 · 2020
Closest in time.
Temporally Grounding Language Queries in Videos by Contextual Boundary-Aware Prediction
Wang, J.; Ma, L.; and Jiang, W. 2020 · 2020
Closest in time.
BERT Representations for Video Question Answering
Yang, Z.; Garcia, N.; Chu, C.; Otani, M.; Nakashima, Y.; and Takemura, H. 2020 · 2020
Closest in time.
Dense regression network for video grounding
Zeng, R.; Xu, H.; Huang, W.; Chen, P.; Tan, M.; and Gan, C. 2020 · 2020
Closest in time.