Fetching the paper…
Reading the bibliography…
In this work, we introduce a novel task - Humancentric Spatio-Temporal Video Grounding (HC-STVG).
J. Duchi, E. Hazan, and Y. Singer, “Adaptive subgradient methods for online learning and stochastic optimization.” Journal of machine learning research , vol. 12, no. 7, 2011
2011
Earlier work this paper cites.
M. Regneri, M. Rohrbach, D. Wetzel, S. Thater, B. Schiele, and M. Pinkal, “Grounding action descriptions in videos,” Trans. Assoc. Comput. Linguistics , vol. 1, pp. 25–36, 2013
2013
Earlier work this paper cites.
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg, “Referitgame: Referring to objects in photographs of natural scenes,” in Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP) , 2014, pp. 787–798
2014
Earlier work this paper cites.
A. Rohrbach, M. Rohrbach, W. Qiu, A. Friedrich, M. Pinkal, and B. Schiele, “Coherent multi-sentence video description with variable level of detail,” in German conference on pattern recognition . Springer, 2014, pp. 184–195
2014
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” in Advances in neural information processing systems , 2015, pp. 91–99
2015
Earlier work this paper cites.
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg, “Modeling context in referring expressions,” in European Conference on Computer Vision . Springer, 2016, pp. 69–85
2016
Earlier work this paper cites.
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy, “Generation and comprehension of unambiguous object descriptions,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 11–20
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
J. Gao, C. Sun, Z. Yang, and R. Nevatia, “Tall: Temporal activity localization via language query,” in Proceedings of the IEEE International Conference on Computer Vision , 2017, pp. 5267–5275
2017
Earlier work this paper cites.
L. Anne Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell, “Localizing moments in video with natural language,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 5803–5812
2017
Earlier work this paper cites.
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. Carlos Niebles, “Dense-captioning events in videos,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 706–715
2017
Earlier work this paper cites.
V. Kalogeiton, P. Weinzaepfel, V. Ferrari, and C. Schmid, “Action tubelet detector for spatio-temporal action localization,” in Proceedings of the IEEE International Conference on Computer Vision , 2017, pp. 4405–4413
2017
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma et al. , “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” International journal of computer vision , vol. 123, no. 1, pp. 32–73, 2017
2017
Earlier work this paper cites.
L. Yu, H. Tan, M. Bansal, and T. L. Berg, “A joint speaker-listener-reinforcer model for referring expressions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 7282–7290
2017
Earlier work this paper cites.
M. Yamaguchi, K. Saito, Y. Ushiku, and T. Harada, “Spatio-temporal person retrieval via natural language queries,” in Proceedings of the IEEE International Conference on Computer Vision , 2017, pp. 1453–1462
2017
Earlier work this paper cites.
B. Li, J. Yan, W. Wu, Z. Zhu, and X. Hu, “High performance visual tracking with siamese region proposal network,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 8971–8980
2018
Earlier work this paper cites.
L. Yu, Z. Lin, X. Shen, J. Yang, X. Lu, M. Bansal, and T. L. Berg, “Mattnet: Modular attention network for referring expression comprehension,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 1307–1315
2018
Cited alongside, same era.
M. Liu, X. Wang, L. Nie, X. He, B. Chen, and T.-S. Chua, “Attentive moment retrieval in videos,” in The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval , 2018, pp. 15–24
2018
Cited alongside, same era.
M. Liu, X. Wang, L. Nie, Q. Tian, B. Chen, and T.-S. Chua, “Cross-modal moment localization in videos,” in Proceedings of the 26th ACM international conference on Multimedia , 2018, pp. 843–851
2018
Cited alongside, same era.
J. Chen, X. Chen, L. Ma, Z. Jie, and T.-S. Chua, “Temporally grounding natural sentence in video,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , 2018, pp. 162–171
2018
Cited alongside, same era.
H. Xu, K. He, B. A. Plummer, L. Sigal, S. Sclaroff, and K. Saenko, “Multilevel language and vision integration for text-to-clip retrieval,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 33, 2019, pp. 9062–9069
2019
Later among the works it cites.
S. Chen and Y.-G. Jiang, “Semantic proposal for activity localization in videos via sentence query,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 33, 2019, pp. 8199–8206
2019
Later among the works it cites.
R. Ge, J. Gao, K. Chen, and R. Nevatia, “Mac: Mining activity concepts for language-based temporal localization,” in 2019 IEEE Winter Conference on Applications of Computer Vision (WACV) . IEEE, 2019, pp. 245–253
2019
Later among the works it cites.
D. Zhang, X. Dai, X. Wang, Y.-F. Wang, and L. S. Davis, “Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 1247–1257
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Gu, C. Sun, D. A. Ross, C. Vondrick, C. Pantofaru, Y. Li, S. Vijayanarasimhan, G. Toderici, S. Ricco, R. Sukthankar, C. Schmid, and J. Malik, “AVA: A video dataset of spatio-temporally localized atomic visual actions,” in IEEE Conference on Computer Vision and Pattern Recognition . IEEE Computer Society, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2018, pp. 2556–2565
2018
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
C. Alberti, J. Ling, M. Collins, and D. Reitter, “Fusion of detected objects in text for visual question answering,” arXiv: Computation and Language , 2019
2019
Cited alongside, same era.
Y. Chen, L. Li, L. Yu, A. E. Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu, “Uniter: Learning universal image-text representations,” arXiv: Computer Vision and Pattern Recognition , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Later among the works it cites.
Y. Yuan, L. Ma, and W. Zhu, “Sentence specified dynamic video thumbnail generation,” in Proceedings of the 27th ACM International Conference on Multimedia , 2019, pp. 2332–2340
2019
Later among the works it cites.
X. Shang, D. Di, J. Xiao, Y. Cao, and T. S. Chua, “Annotating objects and relations in user-generated videos,” in the 2019 , 2019
2019
Later among the works it cites.
C. Lu, L. Chen, C. Tan, X. Li, and J. Xiao, “Debug: A dense bottom-up grounding approach for natural language video localization,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , 2019, pp. 5147–5156
2019
Later among the works it cites.
J. Chen, L. Ma, X. Chen, Z. Jie, and J. Luo, “Localizing natural language in videos,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 33, 2019, pp. 8175–8182
2019
Later among the works it cites.
2019
Later among the works it cites.
S. Huang, S. Liu, T. Hui, J. Han, B. Li, J. Feng, and S. Yan, “Ordnet: Capturing omni-range dependencies for scene parsing,” IEEE Transactions on Image Processing , vol. 29, pp. 8251–8263, 2020
2020
Closest in time.
S. Huang, T. Hui, S. Liu, G. Li, Y. Wei, J. Han, L. Liu, and B. Li, “Referring image segmentation via cross-modal progressive comprehension,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 10 488–10 497
2020
Closest in time.
2020
Closest in time.
Y. Liao, S. Liu, F. Wang, Y. Chen, C. Qian, and J. Feng, “Ppdm: Parallel point detection and matching for real-time human-object interaction detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , June 2020
2020
Closest in time.
G. Li, N. Duan, Y. Fang, M. Gong, D. Jiang, and M. Zhou, “Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training.” in AAAI , 2020, pp. 11 336–11 344
2020
Closest in time.
Z. Zhang, Z. Zhao, Y. Zhao, Q. Wang, H. Liu, and L. Gao, “Where does it exist: Spatio-temporal video grounding for multi-form sentences,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2020
2020
Closest in time.