Fetching the paper…
Reading the bibliography…
We address the problem of language-based temporal localization in untrimmed videos.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Script data for attribute-based recognition of composite activities
M. Rohrbach, M. Regneri, M. Andriluka, S. Amin, M. Pinkal, and B. Schiele · 2012
Earlier work this paper cites.
Grounding action descriptions in videos
M. Regneri, M. Rohrbach, D. Wetzel, S. Thater, B. Schiele, and M. Pinkal · 2013
Earlier work this paper cites.
Active: Activity concept transitions in video event classification
C. Sun and R. Nevatia · 2013
Earlier work this paper cites.
The 2014 sesame multimedia event detection and recounting system
R. Bolles, B. Burns, J. Herson, G. Myers, J. van Hout, W. Wang, J. Wong, E. Yeh, A. Habibian, D. Koelma, et al · 2014
Earlier work this paper cites.
A fast and accurate dependency parser using neural networks
D. Chen and C. Manning · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. Manning · 2014
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Skip-thought vectors
R. Kiros, Y. Zhu, R. R. Salakhutdinov, R. Zemel, R. Urtasun, A. Torralba, and S. Fidler · 2015
Earlier work this paper cites.
Automatic concept discovery from parallel text and visual corpora
C. Sun, C. Gan, and R. Nevatia · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Cited alongside, same era.
Acd: Action concept discovery from image-sentence corpora
J. Gao, C. Sun, and R. Nevatia · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Natural language object retrieval
R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell · 2016
Cited alongside, same era.
Temporal action localization in untrimmed videos via multi-stage CNNs
Z. Shou, D. Wang, and S.-F. Chang · 2016
Cited alongside, same era.
Hollywood in homes: Crowdsourcing data collection for activity understanding
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta · 2016
Cited alongside, same era.
The kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al · 2017
Later among the works it cites.
Cdc: Convolutional-de-convolutional networks for precise temporal action localization in untrimmed videos
Z. Shou, J. Chan, A. Zareian, K. Miyazawa, and S.-F. Chang · 2017
Later among the works it cites.
R-c3d: Region convolutional 3d network for temporal activity detection
H. Xu, A. Das, and K. Saenko · 2017
Later among the works it cites.
Temporal action detection with structured segment networks
Y. Zhao, Y. Xiong, L. Wang, Z. Wu, X. Tang, and D. Lin · 2017
Later among the works it cites.
Temporal action detection with structured segment networks
Y. Zhao, Y. Xiong, L. Wang, Z. Wu, X. Tang, and D. Lin · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Localizing moments in video with natural language
L. Anne Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell · 2017
Cited alongside, same era.
Msrc: Multimodal spatial regression with semantic context for phrase grounding
K. Chen, R. Kovvuri, J. Gao, and R. Nevatia · 2017
Cited alongside, same era.
Query-guided regression network with context policy for phrase grounding
K. Chen, R. Kovvuri, and R. Nevatia · 2017
Cited alongside, same era.
TALL: Temporal activity localization via language query
J. Gao, C. Sun, Z. Yang, and R. Nevatia · 2017
Cited alongside, same era.
TURN TAP: Temporal unit regression network for temporal action proposals
J. Gao, Z. Yang, K. Chen, C. Sun, and R. Nevatia · 2017
Cited alongside, same era.
Cascaded boundary regression for temporal action detection
J. Gao, Z. Yang, and R. Nevatia · 2017
Cited alongside, same era.
Rethinking the faster r-cnn architecture for temporal action localization
Y.-W. Chao, S. Vijayanarasimhan, B. Seybold, D. A. Ross, J. Deng, and R. Sukthankar · 2018
Closest in time.
Knowledge aided consistency for weakly supervised phrase grounding
K. Chen, J. Gao, and R. Nevatia · 2018
Closest in time.
CTAP: Complementary temporal action proposal generation
J. Gao, K. Chen, and R. Nevatia · 2018
Closest in time.
Attentive moment retrieval in videos
M. Liu, X. Wang, L. Nie, X. He, B. Chen, and T.-S. Chua · 2018
Closest in time.
A closer look at spatiotemporal convolutions for action recognition
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri · 2018
Closest in time.
Multi-modal circulant fusion for video-to-language and backward
A. Wu and Y. Han · 2018
Closest in time.