Fetching the paper…
Reading the bibliography…
Temporal grounding aims to localize a video moment which is semantically aligned with a given natural language query.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V.; Debut, L.; Chaumond, J.; and Wolf, T. 2019 · 1910
Earlier work this paper cites.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing
Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; and Brew, J. 2019 · 1910
Earlier work this paper cites.
Long Short-Term Memory
Hochreiter, S.; and Schmidhuber, J. 1997 · 1997
Earlier work this paper cites.
Learning Spatiotemporal Features via Video and Text Pair Discrimination
Li, T.; and Wang, L. 2020 · 2001
Earlier work this paper cites.
Dimensionality Reduction by Learning an Invariant Mapping
Hadsell, R.; Chopra, S.; and LeCun, Y. 2006 · 2006
Earlier work this paper cites.
Visualizing Data using t-SNE
van der Maaten, L.; and Hinton, G. 2008 · 2008
Earlier work this paper cites.
A Simple Yet Effective Method for Video Temporal Grounding with Cross-Modality Attention
Zhang, B.; Li, Y.; Yuan, C.; Xu, D.; Jiang, P.; and Shan, Y. 2020a · 2009
Earlier work this paper cites.
Human-centric Spatio-Temporal Video Grounding With Visual Transformers
Tang, Z.; Liao, Y.; Liu, S.; Li, G.; Jin, X.; Jiang, H.; Yu, Q.; and Xu, D. 2020 · 2011
Earlier work this paper cites.
Script Data for Attribute-Based Recognition of Composite Activities
Rohrbach, M.; Regneri, M.; Andriluka, M.; Amin, S.; Pinkal, M.; and Schiele, B. 2012 · 2012
Earlier work this paper cites.
Multi-Scale 2D Temporal Adjacent Networks for Moment Localization with Natural Language
Zhang, S.; Peng, H.; Fu, J.; Lu, Y.; and Luo, J. 2020c · 2012
Earlier work this paper cites.
Grounding Action Descriptions in Videos
Regneri, M.; Rohrbach, M.; Wetzel, D.; Thater, S.; Schiele, B.; and Pinkal, M. 2013 · 2013
Earlier work this paper cites.
Glove: Global Vectors for Word Representation
Pennington, J.; Socher, R.; and Manning, C. D. 2014 · 2014
Earlier work this paper cites.
ActivityNet: A large-scale video benchmark for human activity understanding
Heilbron, F. C.; Escorcia, V.; Ghanem, B.; and Niebles, J. C. 2015 · 2015
Earlier work this paper cites.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Ren, S.; He, K.; Girshick, R. B.; and Sun, J. 2015 · 2015
Earlier work this paper cites.
Ba, L. J.; Kiros, J. R.; and Hinton, G. E. 2016 · 2016
Earlier work this paper cites.
Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding
Sigurdsson, G. A.; Varol, G.; Wang, X.; Farhadi, A.; Laptev, I.; and Gupta, A. 2016 · 2016
Earlier work this paper cites.
Learning Deep Structure-Preserving Image-Text Embeddings
Wang, L.; Li, Y.; and Lazebnik, S. 2016 · 2016
Earlier work this paper cites.
Temporal Segment Networks: Towards Good Practices for Deep Action Recognition
Wang, L.; Xiong, Y.; Wang, Z.; Qiao, Y.; Lin, D.; Tang, X.; and Gool, L. V. 2016 · 2016
Earlier work this paper cites.
Localizing Moments in Video with Natural Language
Hendricks, L. A.; Wang, O.; Shechtman, E.; Sivic, J.; Darrell, T.; and Russell, B. C. 2017 · 2017
Earlier work this paper cites.
Action Tubelet Detector for Spatio-Temporal Action Localization
Kalogeiton, V.; Weinzaepfel, P.; Ferrari, V.; and Schmid, C. 2017 · 2017
Earlier work this paper cites.
Dense-Captioning Events in Videos
Krishna, R.; Hata, K.; Ren, F.; Fei-Fei, L.; and Niebles, J. C. 2017 · 2017
Earlier work this paper cites.
Top-Down Visual Saliency Guided by Captions
Ramanishka, V.; Das, A.; Zhang, J.; and Saenko, K. 2017 · 2017
Earlier work this paper cites.
Temporal Action Detection with Structured Segment Networks
Zhao, Y.; Xiong, Y.; Wang, L.; Wu, Z.; Tang, X.; and Lin, D. 2017 · 2017
Earlier work this paper cites.
Temporally Grounding Natural Sentence in Video
Chen, J.; Chen, X.; Ma, L.; Jie, Z.; and Chua, T. 2018 · 2018
Cited alongside, same era.
AVA: A Video Dataset of Spatio-Temporally Localized Atomic Visual Actions
Gu, C.; Sun, C.; Ross, D. A.; Vondrick, C.; Pantofaru, C.; Li, Y.; Vijayanarasimhan, S.; Toderici, G.; Ricco, S.; Sukthankar, R.; Schmid, C.; and Malik, J. 2018 · 2018
Cited alongside, same era.
Temporal Modular Networks for Retrieving Complex Compositional Activities in Videos
Liu, B.; Yeung, S.; Chou, E.; Huang, D.; Fei-Fei, L.; and Niebles, J. C. 2018 · 2018
Cited alongside, same era.
Appearance-and-Relation Networks for Video Classification
Wang, L.; Li, W.; Li, W.; and Gool, L. V. 2018 · 2018
Cited alongside, same era.
Unsupervised Feature Learning via Non-Parametric Instance Discrimination
Wu, Z.; Xiong, Y.; Yu, S. X.; and Lin, D. 2018 · 2018
Cited alongside, same era.
Semantic Proposal for Activity Localization in Videos via Sentence Query
Multi-modal Transformer for Video Retrieval
Gabeur, V.; Sun, C.; Alahari, K.; and Schmid, C. 2020 · 2020
Later among the works it cites.
Momentum Contrast for Unsupervised Visual Representation Learning
He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. B. 2020 · 2020
Later among the works it cites.
Supervised Contrastive Learning
Khosla, P.; Teterwak, P.; Wang, C.; Sarna, A.; Tian, Y.; Isola, P.; Maschinot, A.; Liu, C.; and Krishnan, D. 2020 · 2020
Later among the works it cites.
Actions as Moving Points
Li, Y.; Wang, Z.; Wang, L.; and Wu, G. 2020 · 2020
Later among the works it cites.
Moment Retrieval via Cross-Modal Interaction Networks With Query Reconstruction
Lin, Z.; Zhao, Z.; Zhang, Z.; Zhang, Z.; and Cai, D. 2020 · 2020
Later among the works it cites.
End-to-End Learning of Visual Representations From Uncurated Instructional Videos
Miech, A.; Alayrac, J.; Smaira, L.; Laptev, I.; Sivic, J.; and Zisserman, A. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, S.; and Jiang, Y. 2019 · 2019
Cited alongside, same era.
Weakly-Supervised Spatio-Temporally Grounding Natural Sentence in Video
Chen, Z.; Ma, L.; Luo, W.; and Wong, K. K. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2019 · 2019
Cited alongside, same era.
SlowFast Networks for Video Recognition
Feichtenhofer, C.; Fan, H.; Malik, J.; and He, K. 2019 · 2019
Cited alongside, same era.
MAC: Mining Activity Concepts for Language-Based Temporal Localization
Ge, R.; Gao, J.; Chen, K.; and Nevatia, R. 2019 · 2019
Cited alongside, same era.
ExCL: Extractive Clip Localization Using Natural Language Descriptions
Ghosh, S.; Agarwal, A.; Parekh, Z.; and Hauptmann, A. G. 2019 · 2019
Cited alongside, same era.
Read, Watch, and Move: Reinforcement Learning for Temporally Grounding Natural Language Descriptions in Videos
He, D.; Zhao, X.; Huang, J.; Li, F.; Liu, X.; and Wen, S. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Local-Global Video-Text Interactions for Temporal Grounding
Mun, J.; Cho, M.; and Han, B. 2020 · 2020
Later among the works it cites.
Proposal-free Temporal Moment Localization of a Natural-Language Query in Video using Guided Attention
Opazo, C. R.; Marrese-Taylor, E.; Saleh, F. S.; Li, H.; and Gould, S. 2020 · 2020
Later among the works it cites.
Uncovering Hidden Challenges in Query-Based Video Moment Retrieval
Otani, M.; Nakashima, Y.; Rahtu, E.; and Heikkilä, J. 2020 · 2020
Later among the works it cites.
Temporally Grounding Language Queries in Videos by Contextual Boundary-Aware Prediction
Wang, J.; Ma, L.; and Jiang, W. 2020 · 2020
Later among the works it cites.
Tree-Structured Policy Based Progressive Reinforcement Learning for Temporally Language Grounding in Video
Wu, J.; Li, G.; Liu, S.; and Lin, L. 2020 · 2020
Later among the works it cites.
Dense Regression Network for Video Grounding
Zeng, R.; Xu, H.; Huang, W.; Chen, P.; Tan, M.; and Gan, C. 2020 · 2020
Later among the works it cites.
Support-Set Based Cross-Supervision for Video Grounding
Ding, X.; Wang, N.; Zhang, S.; Cheng, D.; Li, X.; Huang, Z.; Tang, M.; and Gao, X. 2021 · 2021
Closest in time.
Fast Video Moment Retrieval
Gao, J.; and Xu, C. 2021 · 2021
Closest in time.
MDETR - Modulated Detection for End-to-End Multi-Modal Understanding
Kamath, A.; Singh, M.; LeCun, Y.; Misra, I.; Synnaeve, G.; and Carion, N. 2021 · 2021
Closest in time.
Context-Aware Biaffine Localizing Network for Temporal Sentence Grounding
Liu, D.; Qu, X.; Dong, J.; Zhou, P.; Cheng, Y.; Wei, W.; Xu, Z.; and Xie, Y. 2021 · 2021
Closest in time.
Interventional Video Grounding With Dual Contrastive Learning
Nan, G.; Qiao, R.; Xiao, Y.; Liu, J.; Leng, S.; Zhang, H.; and Lu, W. 2021 · 2021
Closest in time.
STVGBert: A Visual-Linguistic Transformer Based Framework for Spatio-Temporal Video Grounding
Su, R.; Yu, Q.; and Xu, D. 2021 · 2021
Closest in time.
Augmented 2D-TAN: A Two-stage Approach for Human-centric Spatio-Temporal Video Grounding
Tan, C.; Lin, Z.; Hu, J.; Li, X.; and Zheng, W. 2021 · 2021
Closest in time.
TDN: Temporal Difference Networks for Efficient Action Recognition
Wang, L.; Tong, Z.; Ji, B.; and Wu, G. 2021 · 2021
Closest in time.
2rd Place Solutions in the HC-STVG track of Person in Context Challenge 2021
Yu, Y.; Wang, X.; Hu, W.; Luo, X.; and Li, C. 2021 · 2021
Closest in time.
Cascaded Prediction Network via Segment Tree for Temporal Video Grounding
Zhao, Y.; Zhao, Z.; Zhang, Z.; and Lin, Z. 2021 · 2021
Closest in time.