Sequence to sequence-video to text
S. Venugopalan, M. Rohrbach, J. Donahue, R. Mooney, T. Darrell, and K. Saenko · 2015
Later among the works it cites.
A multi-scale multiple instance video description network
Original
H. Xu, S. Venugopalan, V. Ramanishka, M. Rohrbach, and K. Saenko · 2015
Later among the works it cites.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. C. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio · 2015
Later among the works it cites.
Jointly modeling deep video and compositional text to bridge vision and language in a unified framework
R. Xu, C. Xiong, W. Chen, and J. J. Corso · 2015
Later among the works it cites.
Unsupervised extraction of video highlights via robust recurrent auto-encoders
H. Yang, B. Wang, S. Lin, D. Wipf, M. Guo, and B. Guo · 2015
Later among the works it cites.
Describing videos by exploiting temporal structure
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville · 2015
Later among the works it cites.
Social lstm: Human trajectory prediction in crowded spaces
A. Alahi, K. Goel, V. Ramanathan, A. Robicquet, L. Fei-Fei, and S. Savarese · 2016
Later among the works it cites.
Discovering event structure in continuous narrative perception and memory
C. Baldassano, J. Chen, A. Zadbood, J. W. Pillow, U. Hasson, and K. A. Norman · 2016
Later among the works it cites.
Fast temporal activity proposals for efficient detection of human actions in untrimmed videos
F. Caba Heilbron, J. C. Niebles, and B. Ghanem · 2016
Later among the works it cites.
Daps: Deep action proposals for action understanding
V. Escorcia, F. C. Heilbron, J. C. Niebles, and B. Ghanem · 2016
Later among the works it cites.
Densecap: Fully convolutional localization networks for dense captioning
J. Johnson, A. Karpathy, and L. Fei-Fei · 2016
Later among the works it cites.
Embracing error to enable rapid crowdsourcing
R. A. Krishna, K. Hata, S. Chen, J. Kravitz, D. A. Shamma, L. Fei-Fei, and M. S. Bernstein · 2016
Later among the works it cites.
Learning joint representations of videos and sentences with web image search
M. Otani, Y. Nakashima, E. Rahtu, J. Heikkilä, and N. Yokoya · 2016
Later among the works it cites.
Jointly modeling embedding and translation to bridge video and language
Y. Pan, T. Mei, T. Yao, H. Li, and Y. Rui · 2016
Later among the works it cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta · 2016
Later among the works it cites.
Msr-vtt: A large video description dataset for bridging video and language
J. Xu, T. Mei, T. Yao, and Y. Rui · 2016
Later among the works it cites.
Dense captioning with joint inference and visual context
Original
L. Yang, K. Tang, J. Yang, and L.-J. Li · 2016
Later among the works it cites.
Highlight detection with pairwise deep ranking for first-person video summarization
T. Yao, T. Mei, and Y. Rui · 2016
Later among the works it cites.
Video paragraph captioning using hierarchical recurrent neural networks
H. Yu, J. Wang, Z. Huang, Y. Yang, and W. Xu · 2016
Later among the works it cites.
A hierarchical approach for generating descriptive image paragraphs
J. Krause, J. Johnson, R. Krishna, and L. Fei-Fei · 2017
Closest in time.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, M. Bernstein, and L. Fei-Fei · 2017
Closest in time.