Fetching the paper…
Reading the bibliography…
Video temporal grounding aims to pinpoint a video segment that matches the query description.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
A general approximation framework for direct optimization of information retrieval measures
T. Qin, T.-Y. Liu, and H. Li · 2010
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta · 2016
Earlier work this paper cites.
Highlight detection with pairwise deep ranking for first-person video summarization
T. Yao, T. Mei, and Y. Rui · 2016
Earlier work this paper cites.
Unitbox: An advanced object detection network
J. Yu, Y. Jiang, Z. Wang, Z. Cao, and T. Huang · 2016
Earlier work this paper cites.
Tall: Temporal activity localization via language query
J. Gao, C. Sun, Z. Yang, and R. Nevatia · 2017
Earlier work this paper cites.
Dense-captioning events in videos
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. Carlos Niebles · 2017
Earlier work this paper cites.
Temporally grounding natural sentence in video
J. Chen, X. Chen, L. Ma, Z. Jie, and T.-S. Chua · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
A joint sequence fusion model for video question answering and retrieval
Y. Yu, J. Kim, and G. Kim · 2018
Earlier work this paper cites.
Slowfast networks for video recognition
C. Feichtenhofer, H. Fan, J. Malik, and K. He · 2019
Earlier work this paper cites.
Less is more: Learning highlight detection from video duration
B. Xiong, Y. Kalantidis, D. Ghadiyaram, and K. Grauman · 2019
Cited alongside, same era.
Semantic conditioned dynamic modulation for temporal sentence grounding in videos
Y. Yuan, L. Ma, J. Wang, W. Liu, and W. Zhu · 2019
Cited alongside, same era.
To find where you talk: Temporal sentence localization in video with attention based location regression
Y. Yuan, T. Mei, and W. Zhu · 2019
Cited alongside, same era.
Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment
D. Zhang, X. Dai, X. Wang, Y.-F. Wang, and L. S. Davis · 2019
Cited alongside, same era.
Cross-modal interaction networks for query-based moment retrieval in videos
Z. Zhang, Z. Lin, Z. Zhao, and Z. Xiao · 2019
Cited alongside, same era.
Location-aware graph convolutional networks for video question answering
Learning 2d temporal adjacent networks for moment localization with natural language
S. Zhang, H. Peng, J. Fu, and J. Luo · 2020
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Later among the works it cites.
Vlg-net: Video-language graph matching network for video grounding
M. Soldan, M. Xu, S. Qu, J. Tegner, and B. Ghanem · 2021
Later among the works it cites.
Structured multi-level interaction network for video moment localization via language query
H. Wang, Z.-J. Zha, L. Li, D. Liu, and J. Luo · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Huang, P. Chen, R. Zeng, Q. Du, M. Tan, and C. Gan · 2020
Cited alongside, same era.
Tvr: A large-scale dataset for video-subtitle moment retrieval
J. Lei, L. Yu, T. L. Berg, and M. Bansal · 2020
Cited alongside, same era.
Local-global video-text interactions for temporal grounding
J. Mun, M. Cho, and B. Han · 2020
Cited alongside, same era.
Fine-grained iterative attention network for temporal language localization in videos
X. Qu, P. Tang, Z. Zou, Y. Cheng, J. Dong, P. Zhou, and Z. Xu · 2020
Cited alongside, same era.
Temporally grounding language queries in videos by contextual boundary-aware prediction
J. Wang, L. Ma, and W. Jiang · 2020
Cited alongside, same era.
Dense regression network for video grounding
R. Zeng, H. Xu, W. Huang, P. Chen, M. Tan, and C. Gan · 2020
Cited alongside, same era.
Span-based localizing network for natural language video localization
H. Zhang, A. Sun, W. Jing, and J. T. Zhou · 2020
Cited alongside, same era.
P. Wu, X. He, M. Tang, Y. Lv, and J. Liu · 2021
Later among the works it cites.
Multi-modal interaction graph convolutional network for temporal language localization in videos
Z. Zhang, X. Han, X. Song, Y. Yan, and L. Nie · 2021
Later among the works it cites.
Progressive localization networks for language-based moment localization
Q. Zheng, J. Dong, X. Qu, X. Yang, S. Ji, and X. Wang · 2021
Later among the works it cites.
Ego4d: Around the world in 3,000 hours of egocentric video
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, et al · 2022
Later among the works it cites.
Cone: An efficient coarse-to-fine alignment framework for long video temporal grounding
Z. Hou, W. Zhong, L. Ji, D. Gao, K. Yan, W.-K. Chan, C.-W. Ngo, Z. Shou, and N. Duan · 2022
Later among the works it cites.
Mad: A scalable dataset for language grounding in videos from movie audio descriptions
M. Soldan, A. Pardo, J. L. Alcázar, F. Caba, C. Zhao, S. Giancola, and B. Ghanem · 2022
Later among the works it cites.