Fetching the paper…
Reading the bibliography…
Dense video captioning is a fine-grained video understanding task that involves two sub-problems: localizing distinct events in a long video stream, and generating captions for the localized events.
Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition
S. Guadarrama, N. Krishnamoorthy, G. Malkarnenkar, S. Venugopalan, R. Mooney, T. Darrell, and K. Saenko · 2013
Earlier work this paper cites.
3D Convolutional Neural Networks for Human Action Recognition
S. Ji, W. Xu, M. Yang, and K. Yu · 2013
Earlier work this paper cites.
Action Recognition with Improved Trajectories
H. Wang and C. Schmid · 2013
Earlier work this paper cites.
Coherent multi-sentence video description with variable level of detail
A. Rohrbach, M. Rohrbach, W. Qiu, A. Friedrich, M. Pinkal, and B. Schiele · 2014
Earlier work this paper cites.
Two-stream Convolutional Networks for Action Recognition in Videos
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Earlier work this paper cites.
From captions to visual concepts and back
H. Fang, S. Gupta, F. Iandola, R. K. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. C. Platt, et al · 2015
Earlier work this paper cites.
Fast R-CNN
R. Girshick · 2015
Earlier work this paper cites.
Hierarchical recurrent neural network for document modeling
R. Lin, S. Liu, M. Yang, M. Li, M. Zhou, and S. Li · 2015
Earlier work this paper cites.
Beyond Short Snippets: Deep Networks for Video Classification
J. Y.-H. Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Earlier work this paper cites.
Sequence to sequence – video to text
S. Venugopalan, M. Rohrbach, J. Donahue, R. Mooney, T. Darrell, and K. Saenko · 2015
Earlier work this paper cites.
Translating videos to natural language using deep recurrent neural networks
S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. Mooney, and K. Saenko · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Cited alongside, same era.
Learning to Track for Spatio-Temporal Action Localization
P. Weinzaepfel, Z. Harchaoui, and C. Schmid · 2015
Cited alongside, same era.
A multi-scale multiple instance video description network
H. Xu, S. Venugopalan, V. Ramanishka, M. Rohrbach, and K. Saenko · 2015
Cited alongside, same era.
Describing videos by exploiting temporal structure
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville · 2015
Cited alongside, same era.
Fast Action Proposals for Human Action Detection and Search
Temporal Action Localization in Untrimmed Videos via Multi-stage CNNs
Z. Shou, D. Wang, and S.-F. Chang · 2016
Later among the works it cites.
A Multi-Stream Bi-Directional Recurrent Neural Network for Fine-Grained Action Detection
B. Singh, T. K. Marks, M. Jones, O. Tuzel, and M. Shao · 2016
Later among the works it cites.
Dense captioning with joint inference and visual context
L. Yang, K. Tang, J. Yang, and L.-J. Li · 2016
Later among the works it cites.
End-to-end Learning of Action Detection from Frame Glimpses in Videos
S. Yeung, O. Russakovsky, G. Mori, and L. Fei-Fei · 2016
Later among the works it cites.
Video paragraph captioning using hierarchical recurrent neural networks
H. Yu, J. Wang, Z. Huang, Y. Yang, and W. Xu · 2016
Later among the works it cites.
Sst: Single-stream temporal action proposals
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Yu and J. Yuan · 2015
Cited alongside, same era.
DAPs: Deep Action Proposals for Action Understanding
V. Escorcia, F. C. Heilbron, J. C. Niebles, and B. Ghanem · 2016
Cited alongside, same era.
Densecap: Fully convolutional localization networks for dense captioning
J. Johnson, A. Karpathy, and L. Fei-Fei · 2016
Cited alongside, same era.
A hierarchical approach for generating descriptive image paragraphs
J. Krause, J. Johnson, R. Krishna, and L. Fei-Fei · 2016
Cited alongside, same era.
Learning Activity Progression in LSTMs for Activity Detection and Early Detection
S. Ma, L. Sigal, and S. Sclaroff · 2016
Cited alongside, same era.
Temporal Activity Detection in Untrimmed Videos with Recurrent Neural Networks
A. Montes, A. Salvador, and X. G. i Nieto · 2016
Cited alongside, same era.
Building end-to-end dialogue systems using generative hierarchical neural network models
I. V. Serban, A. Sordoni, Y. Bengio, A. C. Courville, and J. Pineau · 2016
Cited alongside, same era.
S. Buch, V. Escorcia, C. Shen, B. Ghanem, and J. C. Niebles · 2017
Later among the works it cites.
Tall: Temporal activity localization via language query
J. Gao, C. Sun, Z. Yang, and R. Nevatia · 2017
Later among the works it cites.
Turn tap: Temporal unit regression network for temporal action proposals
J. Gao, Z. Yang, K. Chen, C. Sun, and R. Nevatia · 2017
Later among the works it cites.
Localizing moments in video with natural language
L. A. Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell · 2017
Later among the works it cites.
Dense-captioning events in videos
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. C. Niebles · 2017
Later among the works it cites.
R-c3d: Region convolutional 3d network for temporal activity detection
H. Xu, A. Das, and K. Saenko · 2017
Later among the works it cites.