Fetching the paper…
Reading the bibliography…
Temporal sentence grounding (TSG) aims to locate a specific moment from an untrimmed video with a given natural language query.
Script data for attribute-based recognition of composite activities
Marcus Rohrbach, Michaela Regneri, Mykhaylo Andriluka, Sikandar Amin, Manfred Pinkal, and Bernt Schiele · 2012
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Earlier work this paper cites.
Tall: Temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Earlier work this paper cites.
Temporally grounding natural sentence in video
Jingyuan Chen, Xinpeng Chen, Lin Ma, Zequn Jie, and Tat-Seng Chua · 2018
Earlier work this paper cites.
Weakly supervised dense event captioning in videos
Xuguang Duan, Wenbing Huang, Chuang Gan, Jingdong Wang, Wenwu Zhu, and Junzhou Huang · 2018
Earlier work this paper cites.
Three-dimensional attention-based deep ranking model for video highlight detection
Yifan Jiao, Zhetao Li, Shucheng Huang, Xiaoshan Yang, Bin Liu, and Tianzhu Zhang · 2018
Earlier work this paper cites.
Localizing natural language in videos
Jingyuan Chen, Lin Ma, Xinpeng Chen, Zequn Jie, and Jiebo Luo · 2019
Earlier work this paper cites.
Wslln: Weakly supervised natural language localization networks
Mingfei Gao, Larry S Davis, Richard Socher, and Caiming Xiong · 2019
Earlier work this paper cites.
Mac: Mining activity concepts for language-based temporal localization
Runzhou Ge, Jiyang Gao, Kan Chen, and Ram Nevatia · 2019
Earlier work this paper cites.
Cross-modal video moment retrieval with spatial and language-temporal attention
Bin Jiang, Xin Huang, Chao Yang, and Junsong Yuan · 2019
Earlier work this paper cites.
Weakly supervised video moment retrieval from text queries
Niluthpol Chowdhury Mithun, Sujoy Paul, and Amit K Roy-Chowdhury · 2019
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Cited alongside, same era.
Multilevel language and vision integration for text-to-clip retrieval
Huijuan Xu, Kun He, Bryan A Plummer, Leonid Sigal, Stan Sclaroff, and Kate Saenko · 2019
Cited alongside, same era.
To find where you talk: Temporal sentence localization in video with attention based location regression
Yitian Yuan, Tao Mei, and Wenwu Zhu · 2019
Cited alongside, same era.
Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment
Da Zhang, Xiyang Dai, Xin Wang, Yuan-Fang Wang, and Larry S Davis · 2019
Cited alongside, same era.
Look closer to ground better: Weakly-supervised temporal grounding of sentence in video
Zhenfang Chen, Lin Ma, Wenhan Luo, Peng Tang, and Kwan-Yee K Wong · 2020
Structured multi-level interaction network for video moment localization via language query
Hao Wang, Zheng-Jun Zha, Liang Li, Dong Liu, and Jiebo Luo · 2021
Later among the works it cites.
Self-tuning for data-efficient deep learning
Ximei Wang, Jinghan Gao, Mingsheng Long, and Jianmin Wang · 2021
Later among the works it cites.
Fine-grained semantic alignment network for weakly supervised temporal language grounding
Yuechen Wang, Wengang Zhou, and Houqiang Li · 2021
Later among the works it cites.
Boundary proposal network for two-stage natural language video localization
Shaoning Xiao, Long Chen, Songyang Zhang, Wei Ji, Jian Shao, Lu Ye, and Jun Xiao · 2021
Later among the works it cites.
Local correspondence network for weakly supervised temporal sentence grounding
Wenfei Yang, Tianzhu Zhang, Yongdong Zhang, and Feng Wu · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Weakly-supervised video moment retrieval via semantic completion network
Zhijie Lin, Zhou Zhao, Zhu Zhang, Qi Wang, and Huasheng Liu · 2020
Cited alongside, same era.
Local-global video-text interactions for temporal grounding
Jonghwan Mun, Minsu Cho, and Bohyung Han · 2020
Cited alongside, same era.
Yijun Song, Jingwen Wang, Lin Ma, Zhou Yu, and Jun Yu · 2020
Cited alongside, same era.
Temporally grounding language queries in videos by contextual boundary-aware prediction
Jingwen Wang, Lin Ma, and Wenhao Jiang · 2020
Cited alongside, same era.
Learning 2d temporal adjacent networks for moment localization with natural language
Songyang Zhang, Houwen Peng, Jianlong Fu, and Jiebo Luo · 2020
Cited alongside, same era.
Towards bridging event captioner and sentence localizer for weakly supervised dense event captioning
Shaoxiang Chen and Yu-Gang Jiang · 2021
Cited alongside, same era.
Support-set based cross-supervision for video grounding
Xinpeng Ding, Nannan Wang, Shiwei Zhang, De Cheng, Xiaomeng Li, Ziyuan Huang, Mingqian Tang, and Xinbo Gao · 2021
Cited alongside, same era.
Mingxing Zhang, Yang Yang, Xinghan Chen, Yanli Ji, Xing Xu, Jingjing Li, and Heng Tao Shen · 2021
Later among the works it cites.
Video moment retrieval from text queries via single frame annotation
Ran Cui, Tianwen Qian, Pai Peng, Elena Daskalaki, Jingjing Chen, Xiaowei Guo, Huyang Sun, and Yu-Gang Jiang · 2022
Later among the works it cites.
Siod: Single instance annotated per category per image for object detection
Hanjun Li, Xingjia Pan, Ke Yan, Fan Tang, and Wei-Shi Zheng · 2022
Later among the works it cites.
Negative sample matters: A renaissance of metric learning for temporal grounding
Zhenzhi Wang, Limin Wang, Tao Wu, Tianhao Li, and Gangshan Wu · 2022
Later among the works it cites.
Point-supervised video temporal grounding
Zhe Xu, Kun Wei, Xu Yang, and Cheng Deng · 2022
Later among the works it cites.
Tubedetr: Spatio-temporal video grounding with transformers
Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid · 2022
Later among the works it cites.
Video moment retrieval with cross-modal neural architecture search
Xun Yang, Shanshan Wang, Jian Dong, Jianfeng Dong, Meng Wang, and Tat-Seng Chua · 2022
Later among the works it cites.
Weakly supervised video moment localization with contrastive negative sample mining
Minghang Zheng, Yanjie Huang, Qingchao Chen, and Yang Liu · 2022
Later among the works it cites.
Weakly supervised temporal sentence grounding with gaussian-based contrastive proposal learning
Minghang Zheng, Yanjie Huang, Qingchao Chen, Yuxin Peng, and Yang Liu · 2022
Later among the works it cites.
Constraint and union for partially-supervised temporal sentence grounding
Chen Ju, Haicheng Wang, Jinxiang Liu, Chaofan Ma, Ya Zhang, Peisen Zhao, Jianlong Chang, and Qi Tian · 2023
Closest in time.