Fetching the paper…
Reading the bibliography…
Temporal sentence grounding aims to detect the event timestamps described by the natural language query from given untrimmed videos.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Joint inference of groups, events, human roles in aerial videos
Tianmin Shu, Dan Xie, Brandon Rothrock, Sinisa Todorovic, and Song Chun Zhu · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
What’s the point: Semantic segmentation with point supervision
Amy Bearman, Olga Russakovsky, Vittorio Ferrari, and Li Fei-Fei · 2016
Earlier work this paper cites.
Scribblesup: Scribble-supervised convolutional networks for semantic segmentation
Di Lin, Jifeng Dai, Jiaya Jia, Kaiming He, and Jian Sun · 2016
Earlier work this paper cites.
Spot on: Action localization from pointly-supervised proposals
Pascal Mettes, Jan C Van Gemert, and Cees GM Snoek · 2016
Earlier work this paper cites.
Highlight detection with pairwise deep ranking for first-person video summarization
Ting Yao, Tao Mei, and Yong Rui · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Class rectification hard mining for imbalanced deep learning
Qi Dong, Shaogang Gong, and Xiatian Zhu · 2017
Earlier work this paper cites.
Tall: Temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia · 2017
Earlier work this paper cites.
Training object class detectors with click supervision
Dim P Papadopoulos, Jasper RR Uijlings, Frank Keller, and Vittorio Ferrari · 2017
Earlier work this paper cites.
Temporally grounding natural sentence in video
Jingyuan Chen, Xinpeng Chen, Lin Ma, Zequn Jie, and Tat-Seng Chua · 2018
Earlier work this paper cites.
A flexible model for training action localization with varying levels of supervision
Guilhem Chéron, Jean-Baptiste Alayrac, Ivan Laptev, and Cordelia Schmid · 2018
Earlier work this paper cites.
Where are the blobs: Counting by localization with point supervision
Issam H Laradji, Negar Rostamzadeh, Pedro O Pinheiro, David Vazquez, and Mark Schmidt · 2018
Earlier work this paper cites.
Attentive moment retrieval in videos
Meng Liu, Xiang Wang, Liqiang Nie, Xiangnan He, Baoquan Chen, and Tat-Seng Chua · 2018
Earlier work this paper cites.
Cross-modal moment localization in videos
Meng Liu, Xiang Wang, Liqiang Nie, Qi Tian, Baoquan Chen, and Tat-Seng Chua · 2018
Earlier work this paper cites.
Semantic proposal for activity localization in videos via sentence query
Shaoxiang Chen and Yu-Gang Jiang · 2019
Earlier work this paper cites.
Debug: A dense bottom-up grounding approach for natural language video localization
Chujie Lu, Long Chen, Chilie Tan, Xiaolin Li, and Jun Xiao · 2019
Earlier work this paper cites.
Weakly supervised video moment retrieval from text queries
Niluthpol Chowdhury Mithun, Sujoy Paul, and Amit K Roy-Chowdhury · 2019
Earlier work this paper cites.
Action recognition from single timestamp supervision in untrimmed videos
Davide Moltisanti, Sanja Fidler, and Dima Damen · 2019
Earlier work this paper cites.
3c-net: Category count and center loss for weakly-supervised action localization
Sanath Narayan, Hisham Cholakkal, Fahad Shahbaz Khan, and Ling Shao · 2019
Earlier work this paper cites.
Weakly supervised scene parsing with point-based distance metric learning
Rui Qian, Yunchao Wei, Honghui Shi, Jiachen Li, Jiaying Liu, and Thomas Huang · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
Multilevel language and vision integration for text-to-clip retrieval
Huijuan Xu, Kun He, Bryan A Plummer, Leonid Sigal, Stan Sclaroff, and Kate Saenko · 2019
Earlier work this paper cites.
Segregated temporal assembly recurrent networks for weakly supervised multiple action detection
Yunlu Xu, Chengwei Zhang, Zhanzhan Cheng, Jianwen Xie, Yi Niu, Shiliang Pu, and Fei Wu · 2019
Earlier work this paper cites.
To find where you talk: Temporal sentence localization in video with attention based location regression
Yitian Yuan, Tao Mei, and Wenwu Zhu · 2019
Earlier work this paper cites.
Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment
Da Zhang, Xiyang Dai, Xin Wang, Yuan-Fang Wang, and Larry S Davis · 2019
Cited alongside, same era.
Cross-modal interaction networks for query-based moment retrieval in videos
Zhu Zhang, Zhijie Lin, Zhou Zhao, and Zhenxin Xiao · 2019
Cited alongside, same era.
Many-shot from low-shot: Learning to annotate using mixed supervision for object detection
Carlo Biffi, Steven McDonagh, Philip Torr, Aleš Leonardis, and Sarah Parisot · 2020
Cited alongside, same era.
Unsupervised learning of visual features by contrasting cluster assignments
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin · 2020
Cited alongside, same era.
Rethinking the bottom-up framework for query-based video localization
Long Chen, Chujie Lu, Siliang Tang, Jun Xiao, Dong Zhang, Chilie Tan, and Xiaolin Li · 2020
Cited alongside, same era.
Conquer: Contextual query-aware ranking for video corpus moment retrieval
Zhijian Hou, Chong-Wah Ngo, and Wing Kwong Chan · 2021
Later among the works it cites.
Cross-sentence temporal and semantic relations in video activity localisation
Jiabo Huang, Yang Liu, Shaogang Gong, and Hailin Jin · 2021
Later among the works it cites.
Divide and conquer for single-frame temporal action localization
Chen Ju, Peisen Zhao, Siheng Chen, Ya Zhang, Yanfeng Wang, and Qi Tian · 2021
Later among the works it cites.
Learning action completeness from points for weakly-supervised temporal action localization
Pilhyeon Lee and Hyeran Byun · 2021
Later among the works it cites.
Proposal-free video grounding with contextual pyramid network
Kun Li, Dan Guo, and Meng Wang · 2021
Later among the works it cites.
Adaptive proposal generation network for temporal sentence localization in videos
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning modality interaction for temporal sentence localization and event captioning in videos
Shaoxiang Chen, Wenhao Jiang, Wei Liu, and Yu-Gang Jiang · 2020
Cited alongside, same era.
Hierarchical visual-textual graph for temporal activity localization via language
Shaoxiang Chen and Yu-Gang Jiang · 2020
Cited alongside, same era.
Big self-supervised models are strong semi-supervised learners
Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey E Hinton · 2020
Cited alongside, same era.
Look closer to ground better: Weakly-supervised temporal grounding of sentence in video
Zhenfang Chen, Lin Ma, Wenhan Luo, Peng Tang, and Kwan-Yee K Wong · 2020
Cited alongside, same era.
Self-supervised co-training for video representation learning
Tengda Han, Weidi Xie, and Andrew Zisserman · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Cited alongside, same era.
Chen Ju, Peisen Zhao, Ya Zhang, Yanfeng Wang, and Qi Tian · 2020
Cited alongside, same era.
Daizong Liu, Xiaoye Qu, Jianfeng Dong, and Pan Zhou · 2021
Later among the works it cites.
Context-aware biaffine localizing network for temporal sentence grounding
Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou, Yu Cheng, Wei Wei, Zichuan Xu, and Yulai Xie · 2021
Later among the works it cites.
Progressively guide to attend: An iterative alignment framework for temporal sentence grounding
Daizong Liu, Xiaoye Qu, and Pan Zhou · 2021
Later among the works it cites.
Interventional video grounding with dual contrastive learning
Guoshun Nan, Rui Qiao, Yao Xiao, Jun Liu, Sicong Leng, Hao Zhang, and Wei Lu · 2021
Later among the works it cites.
Weakly supervised temporal adjacent network for language grounding
Yuechen Wang, Jiajun Deng, Wengang Zhou, and Houqiang Li · 2021
Later among the works it cites.
Visual co-occurrence alignment learning for weakly-supervised video moment retrieval
Zheng Wang, Jingjing Chen, and Yu-Gang Jiang · 2021
Later among the works it cites.
Boundary proposal network for two-stage natural language video localization
Shaoning Xiao, Long Chen, Songyang Zhang, Wei Ji, Jian Shao, Lu Ye, and Jun Xiao · 2021
Later among the works it cites.
Background-click supervision for temporal action localization
Le Yang, Junwei Han, Tao Zhao, Tianwei Lin, Dingwen Zhang, and Jianxin Chen · 2021
Later among the works it cites.
Local correspondence network for weakly supervised temporal sentence grounding
Wenfei Yang, Tianzhu Zhang, Yongdong Zhang, and Feng Wu · 2021
Later among the works it cites.
Natural language video localization: A revisit in span-based question answering framework
Hao Zhang, Aixin Sun, Wei Jing, Liangli Zhen, Joey Tianyi Zhou, and Rick Siow Mong Goh · 2021
Later among the works it cites.
Video moment retrieval from text queries via single frame annotation
Ran Cui, Tianwen Qian, Pai Peng, Elena Daskalaki, Jingjing Chen, Xiaowei Guo, Huyang Sun, and Yu-Gang Jiang · 2022
Later among the works it cites.
Query-aware video encoder for video moment retrieval
Jiachang Hao, Haifeng Sun, Pengfei Ren, Jingyu Wang, Qi Qi, and Jianxin Liao · 2022
Later among the works it cites.
Sdn: Semantic decoupling network for temporal language grounding
Xun Jiang, Xing Xu, Jingran Zhang, Fumin Shen, Zuo Cao, and Heng Tao Shen · 2022
Later among the works it cites.
Prompting visual-language models for efficient video understanding
Chen Ju, Tengda Han, Kunhao Zheng, Ya Zhang, and Weidi Xie · 2022
Later among the works it cites.
Adaptive mutual supervision for weakly-supervised temporal action localization
Chen Ju, Peisen Zhao, Siheng Chen, Ya Zhang, Xiaoyun Zhang, and Qi Tian · 2022
Later among the works it cites.
Chen Ju, Kunhao Zheng, Jinxiang Liu, Peisen Zhao, Ya Zhang, Jianlong Chang, Yanfeng Wang, and Qi Tian · 2022
Later among the works it cites.
Compositional temporal grounding with structured variational cross-graph correspondence learning
Juncheng Li, Junlin Xie, Long Qian, Linchao Zhu, Siliang Tang, Fei Wu, Yi Yang, Yueting Zhuang, and Xin Eric Wang · 2022
Later among the works it cites.
Memory-guided semantic learning network for temporal sentence grounding
Daizong Liu, Xiaoye Qu, Xing Di, Yu Cheng, Zichuan Xu, and Pan Zhou · 2022
Later among the works it cites.
Reducing the vision and language bias for temporal sentence grounding
Daizong Liu, Xiaoye Qu, and Wei Hu · 2022
Later among the works it cites.
Exploring motion and appearance information for temporal sentence grounding
Daizong Liu, Xiaoye Qu, Pan Zhou, and Yang Liu · 2022
Later among the works it cites.
Weakly supervised video moment localization with contrastive negative sample mining
Minghang Zheng, Yanjie Huang, Qingchao Chen, and Yang Liu · 2022
Later among the works it cites.
Weakly supervised temporal sentence grounding with gaussian-based contrastive proposal learning
Minghang Zheng, Yanjie Huang, Qingchao Chen, Yuxin Peng, and Yang Liu · 2022
Later among the works it cites.