Fetching the paper…
Reading the bibliography…
Spatial-Temporal Video Grounding (STVG) is a challenging task which aims to localize the spatio-temporal tube of the interested object semantically according to a natural language query.
“Glove: Global vectors for word representation,”
Jeffrey Pennington, Richard Socher, and Christopher D Manning, · 2014
Earlier work this paper cites.
“Empirical evaluation of gated recurrent neural networks on sequence modeling,”
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Natural language object retrieval,”
Ronghang Hu, Huazhe Xu, Marcus Rohrbach, Jiashi Feng, Kate Saenko, and Trevor Darrell, · 2016
Earlier work this paper cites.
“Grounding of textual phrases in images by reconstruction,”
Anna Rohrbach, Marcus Rohrbach, Ronghang Hu, Trevor Darrell, and Bernt Schiele, · 2016
Earlier work this paper cites.
“Tall: Temporal activity localization via language query,”
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia, · 2017
Earlier work this paper cites.
“Localizing moments in video with natural language,”
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell, · 2017
Earlier work this paper cites.
“Spatio-temporal person retrieval via natural language queries,”
Masataka Yamaguchi, Kuniaki Saito, Yoshitaka Ushiku, and Tatsuya Harada, · 2017
Earlier work this paper cites.
“Quo vadis, action recognition? a new model and the kinetics dataset,”
Joao Carreira and Andrew Zisserman, · 2017
Cited alongside, same era.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“Focal loss for dense object detection,”
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár, · 2017
Cited alongside, same era.
“Mattnet: Modular attention network for referring expression comprehension,”
Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg, · 2018
Cited alongside, same era.
“Weakly-supervised video object grounding from text by loss weighting and object interaction,”
Luowei Zhou, Nathan Louis, and Jason J Corso, · 2018
Cited alongside, same era.
“Dynamic graph attention for referring expression comprehension,”
Sibei Yang, Guanbin Li, and Yizhou Yu, · 2019
Later among the works it cites.
“Weakly-supervised spatio-temporally grounding natural sentence in video,”
Zhenfang Chen, Lin Ma, Wenhan Luo, and Kwan-Yee K Wong, · 2019
Later among the works it cites.
“Localizing natural language in videos,”
Jingyuan Chen, Lin Ma, Xinpeng Chen, Zequn Jie, and Jiebo Luo, · 2019
Later among the works it cites.
“Jointly cross-and self-modal graph attention network for query-based moment localization,”
Daizong Liu, Xiaoye Qu, Xiao-Yang Liu, Jianfeng Dong, Pan Zhou, and Zichuan Xu, · 2020
Later among the works it cites.
“Where does it exist: Spatio-temporal video grounding for multi-form sentences,”
Zhu Zhang, Zhou Zhao, Yang Zhao, Qi Wang, Huasheng Liu, and Lianli Gao, · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Arun Balajee Vasudevan, Dengxin Dai, and Luc Van Gool, · 2018
Cited alongside, same era.
“Cross-modal relationship inference for grounding referring expressions,”
Sibei Yang, Guanbin Li, and Yizhou Yu, · 2019
Cited alongside, same era.
“Context-aware biaffine localizing network for temporal sentence grounding,”
Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou, Yu Cheng, Wei Wei, Zichuan Xu, and Yulai Xie, · 2021
Later among the works it cites.