Fetching the paper…
Reading the bibliography…
Temporal Sentence Grounding in Videos (TSGV), i.e., grounding a natural language sentence which indicates complex human activities in a long and untrimmed video sequence, has received unprecedented attentions over the last few years.
Grounding action descriptions in videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal. 2013 · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation. In EMNLP
Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014 · 2014
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks. In ICCV
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. 2015 · 2015
Earlier work this paper cites.
Temporal action localization in untrimmed videos via multi-stage cnns. In CVPR
Zheng Shou, Dongang Wang, and Shih-Fu Chang. 2016 · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding. In ECCV
Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. 2016 · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition. In ECCV
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. 2016 · 2016
Earlier work this paper cites.
Localizing moments in video with natural language. In ICCV
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. 2017 · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset. In CVPR
Joao Carreira and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
Tall: Temporal activity localization via language query. In ICCV
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia. 2017 · 2017
Earlier work this paper cites.
Dense-captioning events in videos. In ICCV
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. 2017 · 2017
Earlier work this paper cites.
Temporally grounding natural sentence in video. In EMNLP
Jingyuan Chen, Xinpeng Chen, Lin Ma, Zequn Jie, and Tat-Seng Chua. 2018 · 2018
Earlier work this paper cites.
Weakly supervised dense event captioning in videos. In NeurIPS
Xuguang Duan, Wenbing Huang, Chuang Gan, Jingdong Wang, Wenwu Zhu, and Junzhou Huang. 2018 · 2018
Cited alongside, same era.
Localizing moments in video with temporal language. In EMNLP
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. 2018 · 2018
Cited alongside, same era.
Val: Visual-attention action localizer. In PCM
Xiaomeng Song and Yahong Han. 2018 · 2018
Cited alongside, same era.
WSLLN: Weakly Supervised Natural Language Localization Networks. In EMNLP
Mingfei Gao, Larry Davis, Richard Socher, and Caiming Xiong. 2019 · 2019
Cited alongside, same era.
Mac: Mining activity concepts for language-based temporal localization. In WACV
Runzhou Ge, Jiyang Gao, Kan Chen, and Ram Nevatia. 2019 · 2019
Cited alongside, same era.
Tripping through time: Efficient localization of activities in videos. In arXiv
wman: Weakly-supervised moment alignment network for text-based video segment retrieval. In arXiv
Reuben Tan, Huijuan Xu, Kate Saenko, and Bryan A Plummer. 2019 · 2019
Later among the works it cites.
Language-driven temporal activity localization: A semantic matching reinforcement learning model. In CVPR
Weining Wang, Yan Huang, and Liang Wang. 2019 · 2019
Later among the works it cites.
Multilevel language and vision integration for text-to-clip retrieval. In AAAI
Huijuan Xu, Kun He, Bryan A Plummer, Leonid Sigal, Stan Sclaroff, and Kate Saenko. 2019 · 2019
Later among the works it cites.
Rethinking the Bottom-Up Framework for Query-Based Video Localization.. In AAAI
Long Chen, Chujie Lu, Siliang Tang, Jun Xiao, Dong Zhang, Chilie Tan, and Xiaolin Li. 2020 · 2020
Later among the works it cites.
Uncovering Hidden Challenges in Query-Based Video Moment Retrieval. In BMVC
Mayu Otani, Yuta Nakashima, Esa Rahtu, and Janne Heikkilä. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Meera Hahn, Asim Kadav, James M Rehg, and Hans Peter Graf. 2019 · 2019
Cited alongside, same era.
Read, watch, and move: Reinforcement learning for temporally grounding natural language descriptions in videos. In AAAI
Dongliang He, Xiang Zhao, Jizhou Huang, Fu Li, Xiao Liu, and Shilei Wen. 2019 · 2019
Cited alongside, same era.
Cross-modal video moment retrieval with spatial and language-temporal attention. In ICMR
Bin Jiang, Xin Huang, Chao Yang, and Junsong Yuan. 2019 · 2019
Cited alongside, same era.
Debug: A dense bottom-up grounding approach for natural language video localization. In EMNLP
Chujie Lu, Long Chen, Chilie Tan, Xiaolin Li, and Jun Xiao. 2019 · 2019
Cited alongside, same era.
Weakly supervised video moment retrieval from text queries. In CVPR
Niluthpol Chowdhury Mithun, Sujoy Paul, and Amit K Roy-Chowdhury. 2019 · 2019
Cited alongside, same era.
Attentive moment retrieval in videos. In SIGIR
Meng Liu, Xiang Wang, Liqiang Nie, Xiangnan He, Baoquan Chen, and Tat-Seng Chua. 2018a
Cited in the paper.
Cross-modal moment localization in videos. In ACM MM
Meng Liu, Xiang Wang, Liqiang Nie, Qi Tian, Baoquan Chen, and Tat-Seng Chua. 2018b
Cited in the paper.
Weakly-supervised multi-level attentional reconstruction network for grounding textual queries in videos. In arXiv
Yijun Song, Jingwen Wang, Lin Ma, Zhou Yu, and Jun Yu. 2020 · 2020
Later among the works it cites.
Tree-Structured Policy based Progressive Reinforcement Learning for Temporally Language Grounding in Video. In AAAI
Jie Wu, Guanbin Li, Si Liu, and Liang Lin. 2020 · 2020
Later among the works it cites.
Dense regression network for video grounding. In CVPR
Runhao Zeng, Haoming Xu, Wenbing Huang, Peihao Chen, Mingkui Tan, and Chuang Gan. 2020 · 2020
Later among the works it cites.
Learning 2D Temporal Adjacent Networks forMoment Localization with Natural Language. In AAAI
Songyang Zhang, Houwen Peng, Jianlong Fu, and Jiebo Luo. 2020 · 2020
Later among the works it cites.
Boundary Proposal Network for Two-Stage Natural Language Video Localization. In AAAI
Shaoning Xiao, Long Chen, Songyang Zhang, Wei Ji, Jian Shao, Lu Ye, and Jun Xiao. 2021 · 2021
Closest in time.