Fetching the paper…
Reading the bibliography…
In this paper, we propose an end-to-end capsule network for pixel level localization of actors and actions present in a video.
Transforming auto-encoders
G. E. Hinton, A. Krizhevsky, and S. D. Wang · 2011
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Towards understanding action recognition
H. Jhuang, J. Gall, S. Zuffi, C. Schmid, and M. J. Black · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Can humans fly? action understanding with multiple classes of actors
C. Xu, S.-H. Hsieh, C. Xiong, and J. J. Corso · 2015
Earlier work this paper cites.
Segmentation from natural language expressions
R. Hu, M. Rohrbach, and T. Darrell · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta · 2016
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Cited alongside, same era.
Tall: Temporal activity localization via language query
J. Gao, C. Sun, Z. Yang, and R. Nevatia · 2017
Cited alongside, same era.
Localizing moments in video with natural language
L. A. Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell · 2017
Cited alongside, same era.
Going deeper into action recognition: A survey
S. Herath, M. Harandi, and F. Porikli · 2017
Cited alongside, same era.
Tube convolutional neural network (t-cnn) for action detection in videos
Dynamic routing between capsules
S. Sabour, N. Frosst, and G. E. Hinton · 2017
Later among the works it cites.
Spatio-temporal person retrieval via natural language queries
M. Yamaguchi, K. Saito, Y. Ushiku, and T. Harada · 2017
Later among the works it cites.
Videocapsulenet: A simplified network for action detection
K. Duarte, Y. S. Rawat, and M. Shah · 2018
Closest in time.
Actor and action video segmentation from a sentence
K. Gavrilyuk, A. Ghodrati, Z. Li, and C. G. Snoek · 2018
Closest in time.
Ava: A video dataset of spatio-temporally localized atomic visual actions
C. Gu, C. Sun, S. Vijayanarasimhan, C. Pantofaru, D. A. Ross, G. Toderici, Y. Li, S. Ricco, R. Sukthankar, C. Schmid, et al · 2018
Closest in time.
Matrix capsules with em routing
G. Hinton, S. Sabour, and N. Frosst · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Hou, C. Chen, and M. Shah · 2017
Cited alongside, same era.
Action tubelet detector for spatio-temporal action localization
V. Kalogeiton, P. Weinzaepfel, V. Ferrari, and C. Schmid · 2017
Cited alongside, same era.
The kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al · 2017
Cited alongside, same era.
Tracking by natural language specification
Z. Li, R. Tao, E. Gavves, C. G. Snoek, A. W. Smeulders, et al · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer · 2017
Cited alongside, same era.
Person search with natural language description
S. Li, T. Xiao, H. Li, B. Zhou, D. Yue, and X. Wang
Cited in the paper.
Closest in time.
End-to-end joint semantic segmentation of actors and actions in video
J. Ji, S. Buch, A. Soto, and J. C. Niebles · 2018
Closest in time.
Capsules for object segmentation
R. LaLonde and U. Bagci · 2018
Closest in time.
Key-word-aware network for referring expression image segmentation
H. Shi, H. Li, F. Meng, and Q. Wu · 2018
Closest in time.
Investigating capsule networks with dynamic routing for text classification
W. Zhao, J. Ye, M. Yang, Z. Lei, S. Zhang, and Z. Zhao · 2018
Closest in time.