Fetching the paper…
Reading the bibliography…
This paper targets the task of language-based video moment localization.
Weakly-supervised video object grounding by exploring spatio-temporal contexts. In Proceedings of the 28th ACM international conference on multimedia . 1939–1947
Xun Yang, Xueliang Liu, Meng Jian, Xinjian Gao, and Meng Wang. 2020b · 1947
Earlier work this paper cites.
Script data for attribute-based recognition of composite activities. In European Conference on Computer Vision . 144–157
Marcus Rohrbach, Michaela Regneri, Mykhaylo Andriluka, Sikandar Amin, Manfred Pinkal, and Bernt Schiele. 2012 · 2012
Earlier work this paper cites.
Near-lossless semantic video summarization and its applications to video analysis
Tao Mei, Lin-Xie Tang, Jinhui Tang, and Xian-Sheng Hua. 2013 · 2013
Earlier work this paper cites.
Grounding action descriptions in videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal. 2013 · 2013
Earlier work this paper cites.
GloVe: Global Vectors for Word Representation. In Proceedings of the Conference on Empirical Methods in Natural Language Processing . 1532–1543
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 961–970
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization. In International Conference for Learning Representations . 1–15
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in Neural Information Processing Systems . 91–99
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition. In International Conference for Learning Representations . 1–14
Karen Simonyan and Andrew Zisserman. 2015 · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks. In Proceedings of the IEEE International Conference on Computer Vision . 4489–4497
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. 2015 · 2015
Earlier work this paper cites.
Temporal action localization in untrimmed videos via multi-stage cnns. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 1049–1058
Zheng Shou, Dongang Wang, and Shih-Fu Chang. 2016 · 2016
Earlier work this paper cites.
Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding. In European Conference on Computer Vision . 510–526
Gunnar A. Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. 2016 · 2016
Earlier work this paper cites.
Semantic feature mining for video event understanding
Xiaoshan Yang, Tianzhu Zhang, and Changsheng Xu. 2016 · 2016
Earlier work this paper cites.
SST: Single-Stream Temporal Action Proposals. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 2911–2920
Shyamal Buch, Victor Escorcia, Chuanqi Shen, Bernard Ghanem, and Juan Carlos Niebles. 2017 · 2017
Earlier work this paper cites.
Localizing Moments in Video With Natural Language. In Proceedings of the IEEE International Conference on Computer Vision . 5803–5812
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. 2017 · 2017
Earlier work this paper cites.
Dense-Captioning Events in Videos. In Proceedings of the IEEE International Conference on Computer Vision . 706–715
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. 2017 · 2017
Earlier work this paper cites.
Single shot temporal action detection. In Proceedings of the ACM international Conference on Multimedia . 988–996
Tianwei Lin, Xu Zhao, and Zheng Shou. 2017 · 2017
Earlier work this paper cites.
Cdc: Convolutional-de-convolutional networks for precise temporal action localization in untrimmed videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 5734–5743
Zheng Shou, Jonathan Chan, Alireza Zareian, Kazuyuki Miyazawa, and Shih-Fu Chang. 2017 · 2017
Earlier work this paper cites.
Attention is all you need. In Advances in Neural Information Processing Systems . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Rethinking the faster r-cnn architecture for temporal action localization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 1130–1139
Yu-Wei Chao, Sudheendra Vijayanarasimhan, Bryan Seybold, David A Ross, Jia Deng, and Rahul Sukthankar. 2018 · 2018
Earlier work this paper cites.
Temporally Grounding Natural Sentence in Video. In Proceedings of the Conference on Empirical Methods in Natural Language Processing . 162–171
Jingyuan Chen, Xinpeng Chen, Lin Ma, Zequn Jie, and Tat-Seng Chua. 2018 · 2018
Earlier work this paper cites.
Cross-media similarity evaluation for web image retrieval in the wild
Jianfeng Dong, Xirong Li, and Duanqing Xu. 2018 · 2018
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 32. 3942–3951
Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville. 2018 · 2018
Cited alongside, same era.
Yolov3: An incremental improvement
Joseph Redmon and Ali Farhadi. 2018 · 2018
Cited alongside, same era.
Relation attention for temporal action localization
Peihao Chen, Chuang Gan, Guangyao Shen, Wenbing Huang, Runhao Zeng, and Mingkui Tan. 2019 · 2019
Cited alongside, same era.
Semantic Proposal for Activity Localization in Videos via Sentence Query. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 33. 8199–8206
Shaoxiang Chen and Yu-Gang Jiang. 2019 · 2019
Cited alongside, same era.
MAC: Mining Activity Concepts for Language-based Temporal Localization. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision . 245–253
Deep neighborhood component analysis for visual similarity modeling
Xueliang Liu, Xun Yang, Meng Wang, and Richang Hong. 2020b · 2020
Later among the works it cites.
Local-Global Video-Text Interactions for Temporal Grounding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 10810–10819
Jonghwan Mun, Minsu Cho, and Bohyung Han. 2020 · 2020
Later among the works it cites.
Proposal-free Temporal Moment Localization of a Natural-Language Query in Video using Guided Attention. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision . 2464–2473
Cristian Rodriguez, Edison Marrese-Taylor, Fatemeh Sadat Saleh, HONGDONG LI, and Stephen Gould. 2020 · 2020
Later among the works it cites.
Temporally Grounding Language Queries in Videos by Contextual Boundary-Aware Prediction.. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 34. 12168–12175
Jingwen Wang, Lin Ma, and Wenhao Jiang. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Runzhou Ge, Jiyang Gao, Kan Chen, and Ram Nevatia. 2019 · 2019
Cited alongside, same era.
Cross-Modal Video Moment Retrieval with Spatial and Language-Temporal Attention. In Proceedings of the International Conference on Multimedia Retrieval . 217–225
Bin Jiang, Xin Huang, Chao Yang, and Junsong Yuan. 2019 · 2019
Cited alongside, same era.
DEBUG: A Dense Bottom-Up Grounding Approach for Natural Language Video Localization. In Proceedings of the Conference on Empirical Methods in Natural Language Processing and the International Joint Conference on Natural Language Processing . 5147–5156
Chujie Lu, Long Chen, Chilie Tan, Xiaolin Li, and Jun Xiao. 2019 · 2019
Cited alongside, same era.
Weakly supervised video moment retrieval from text queries. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 11592–11601
Niluthpol Chowdhury Mithun, Sujoy Paul, and Amit K Roy-Chowdhury. 2019 · 2019
Cited alongside, same era.
CM-GANs: Cross-modal generative adversarial networks for common representation learning
Yuxin Peng and Jinwei Qi. 2019 · 2019
Cited alongside, same era.
Annotating objects and relations in user-generated videos. In Proceedings of the 2019 on International Conference on Multimedia Retrieval . 279–287
Xindi Shang, Donglin Di, Junbin Xiao, Yu Cao, Xun Yang, and Tat-Seng Chua. 2019 · 2019
Cited alongside, same era.
Fcos: Fully convolutional one-stage object detection. In Proceedings of the IEEE International Conference on Computer Vision . 9627–9636
Zhi Tian, Chunhua Shen, Hao Chen, and Tong He. 2019 · 2019
Cited alongside, same era.
Language-driven Temporal Activity Localization: A Semantic Matching Reinforcement Learning Model. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 334–343
Weining Wang, Yan Huang, and Liang Wang. 2019 · 2019
Cited alongside, same era.
Junbin Xiao, Xindi Shang, Xun Yang, Sheng Tang, and Tat-Seng Chua. 2020 · 2020
Later among the works it cites.
Semantic Conditioned Dynamic Modulation for Temporal Sentence Grounding in Videos
Yitian Yuan, Lin Ma, Jingwen Wang, Wei Liu, and Wenwu Zhu. 2020 · 2020
Later among the works it cites.
Dense regression network for video grounding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 10287–10296
Runhao Zeng, Haoming Xu, Wenbing Huang, Peihao Chen, Mingkui Tan, and Chuang Gan. 2020 · 2020
Later among the works it cites.
Temporal Textual Localization in Video via Adversarial Bi-Directional Interaction Networks
Zijian Zhang, Zhou Zhao, Zhu Zhang, Zhijie Lin, Qi Wang, and Richang Hong. 2020d · 2020
Later among the works it cites.
Learning Video Moment Retrieval Without a Single Annotated Video
Junyu Gao and Changsheng Xu. 2021b · 2021
Closest in time.
Video Moment Localization via Deep Cross-Modal Hashing
Yupeng Hu, Meng Liu, Xiaobin Su, Zan Gao, and Liqiang Nie. 2021 · 2021
Closest in time.
Interventional video relation detection. In Proceedings of the 29th ACM International Conference on Multimedia . 4091–4099
Yicong Li, Xun Yang, Xindi Shang, and Tat-Seng Chua. 2021 · 2021
Closest in time.
Context-aware Biaffine Localizing Network for Temporal Sentence Grounding
Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou, Yu Cheng, Wei Wei, Zichuan Xu, and Yulai Xie. 2021c · 2021
Closest in time.
Single-shot Semantic Matching Network for Moment Localization in Videos
Xinfang Liu, Xiushan Nie, Junya Teng, Li Lian, and Yilong Yin. 2021b · 2021
Closest in time.
Centerness-Aware Network for Temporal Action Proposal
Yuan Liu, Jingyuan Chen, Xinpeng Chena, Bing Deng, Jianqiang Huang, and Xiansheng Hua. 2021a · 2021
Closest in time.
Interaction-Integrated Network for Natural Language Moment Localization
Ke Ning, Lingxi Xie, Jianzhuang Liu, Fei Wu, and Qi Tian. 2021 · 2021
Closest in time.
MABAN: Multi-Agent Boundary-Aware Network for Natural Language Moment Retrieval
Xiaoyang Sun, Hanli Wang, and Bin He. 2021 · 2021
Closest in time.
Boundary Proposal Network for Two-Stage Natural Language Video Localization
Shaoning Xiao, Long Chen, Songyang Zhang, Wei Ji, Jian Shao, Lu Ye, and Jun Xiao. 2021 · 2021
Closest in time.
Cross-Modal Hybrid Feature Fusion for Image-Sentence Matching
Xing Xu, Yifan Wang, Yixuan He, Yang Yang, Alan Hanjalic, and Heng Tao Shen. 2021 · 2021
Closest in time.
Deconfounded Video Moment Retrieval with Causal Intervention. In Proceedings of the ACM SIGIR Conference on Research and Development in Information Retrieval
Xun Yang, Fuli Feng, Wei Ji, Meng Wang, and Tat-Seng Chua. 2021 · 2021
Closest in time.
Video Moment Retrieval with Cross-Modal Neural Architecture Search
Xun Yang, Shanshan Wang, Jian Dong, Jianfeng Dong, Meng Wang, and Tat-Seng Chua. 2022 · 2022
Closest in time.
Logan: Latent graph co-attention network for weakly-supervised video moment retrieval. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision . 2083–2092
Reuben Tan, Huijuan Xu, Kate Saenko, and Bryan A Plummer. 2021b · 2092
Closest in time.