Fetching the paper…
Reading the bibliography…
This notebook paper presents an overview and comparative analysis of our systems designed for the following five tasks in ActivityNet Challenge 2018: temporal action proposals, temporal action localization, dense-captioning events in videos, trimmed action recognition, and spatio-temporal action localization.
Activitynet: A large-scale video benchmark for human activity understanding
F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. Carlos Niebles · 2015
Earlier work this paper cites.
Msr asia msm at thumos challenge 2015
Z. Qiu, Q. Li, T. Yao, T. Mei, and Y. Rui · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Earlier work this paper cites.
Towards good practices for very deep two-stream convnets
L. Wang, Y. Xiong, Z. Wang, and Y. Qiao · 2015
Earlier work this paper cites.
Compact bilinear pooling
Y. Gao, O. Beijbom, N. Zhang, and T. Darrell · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Action recognition by learning deep multi-granular spatio-temporal video representation
Q. Li, Z. Qiu, T. Yao, T. Mei, Y. Rui, and J. Luo · 2016
Earlier work this paper cites.
Jointly modeling embedding and translation to bridge video and language
Y. Pan, T. Mei, T. Yao, H. Li, and Y. Rui · 2016
Cited alongside, same era.
Msr asia msm at activitynet challenge 2016
Z. Qiu, D. Li, C. Gan, T. Yao, T. Mei, and Y. Rui · 2016
Cited alongside, same era.
Self-critical sequence training for image captioning
S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel · 2016
Cited alongside, same era.
Deformable convolutional networks
J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei · 2017
Cited alongside, same era.
The kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al · 2017
Cited alongside, same era.
Deep quantization: Encoding convolutional activations with deep generative model
Z. Qiu, T. Yao, and T. Mei · 2017
Later among the works it cites.
Learning spatio-temporal representation with pseudo-3d residual networks
Z. Qiu, T. Yao, and T. Mei · 2017
Later among the works it cites.
CDC: Convolutional-De-Convolutional Network for Precise Temporal Action Localization in Untrimmed Videos
Z. Shou, J. Chan, A. Zareian, K. Miyazawa, and S.-F. Chang · 2017
Later among the works it cites.
A Pursuit of Temporal Accuracy in General Activity Detection
Y. Xiong, Y. Zhao, L. Wang, D. Lin, and X. Tang · 2017
Later among the works it cites.
Msr asia msm at activitynet challenge 2017: Trimmed action recognition, temporal action proposals and dense-captioning events in videos
T. Yao, Y. Li, Z. Qiu, F. Long, Y. Pan, D. Li, and T. Mei · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Temporal Convolution Based Action Proposal: Submission to ActivityNet 2017
T. Lin, X. Zhao, and Z. Shou · 2017
Cited alongside, same era.
Seeing bot
Y. Pan, Z. Qiu, T. Yao, H. Li, and T. Mei · 2017
Cited alongside, same era.
To create what you tell: Generating videos from captions
Y. Pan, Z. Qiu, T. Yao, H. Li, and T. Mei · 2017
Cited alongside, same era.
Video captioning with transferred semantic attributes
Y. Pan, T. Yao, H. Li, and T. Mei · 2017
Cited alongside, same era.
Incorporating copying mechanism in image captioning for learning novel objects
T. Yao, Y. Pan, Y. Li, and T. Mei · 2017
Later among the works it cites.
Boosting image captioning with attributes
T. Yao, Y. Pan, Y. Li, Z. Qiu, and T. Mei · 2017
Later among the works it cites.
Jointly localizing and describing events for dense video captioning
Y. Li, T. Yao, Y. Pan, H. Chao, and T. Mei · 2018
Closest in time.