Fetching the paper…
Reading the bibliography…
In this paper, we address the problem of spatio-temporal person retrieval from multiple videos using a natural language query, in which we output a tube (i.e., a sequence of bounding boxes) which encloses the person described by the query.
Video google: A text retrieval approach to object matching in videos
J. Sivic, A. Zisserman, et al · 2003
Earlier work this paper cites.
Video retrieval using high level features: Exploiting query matching and confidence-based weighting
S.-Y. Neo, J. Zhao, M.-Y. Kan, and T.-S. Chua · 2006
Earlier work this paper cites.
Video search in concept subspace: a text-like paradigm
X. Li, D. Wang, J. Li, and B. Zhang · 2007
Earlier work this paper cites.
Adding semantics to detectors for video retrieval
C. G. Snoek, B. Huurnink, L. Hollink, M. De Rijke, G. Schreiber, and M. Worring · 2007
Earlier work this paper cites.
A duality based approach for realtime tv-l 1 optical flow
C. Zach, T. Pock, and H. Bischof · 2007
Earlier work this paper cites.
Action mach a spatio-temporal maximum average correlation height filter for action recognition
M. D. Rodriguez, J. Ahmed, and M. Shah · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Collecting image annotations using amazon’s mechanical turk
C. Rashtchian, P. Young, M. Hodosh, and J. Hockenmaier · 2010
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Spatiotemporal deformable part models for action detection
Y. Tian, R. Sukthankar, and M. Shah · 2013
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Earlier work this paper cites.
Action localization with tubelets from motion
M. Jain, J. Van Gemert, H. Jégou, P. Bouthemy, and C. G. Snoek · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Earlier work this paper cites.
Visual semantic search: Retrieving videos via complex textual queries
D. Lin, S. Fidler, C. Kong, and R. Urtasun · 2014
Earlier work this paper cites.
Spatio-temporal object detection proposals
D. Oneata, J. Revaud, J. Verbeek, and C. Schmid · 2014
Earlier work this paper cites.
Seeing what you’re told: Sentence-guided activity recognition in video
N. Siddharth, A. Barbu, and J. Mark Siskind · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Cited alongside, same era.
Video action detection with relational dynamic-poselets
L. Wang, Y. Qiao, and X. Tang · 2014
Cited alongside, same era.
Activitynet: A large-scale video benchmark for human activity understanding
F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. Carlos Niebles · 2015
Cited alongside, same era.
Action detection by implicit intentional motion clustering
W. Chen and J. J. Corso · 2015
Cited alongside, same era.
Finding action tubes
G. Gkioxari and J. Malik · 2015
Cited alongside, same era.
Apt: Action localization proposals from dense trajectories
J. van Gemert, M. Jain, E. Gati, and C. Snoek · 2015
Later among the works it cites.
Towards good practices for very deep two-stream convnets
L. Wang, Y. Xiong, Z. Wang, and Y. Qiao · 2015
Later among the works it cites.
Learning to track for spatio-temporal action localization
P. Weinzaepfel, Z. Harchaoui, and C. Schmid · 2015
Later among the works it cites.
Fast action proposals for human action detection and search
G. Yu and J. Yuan · 2015
Later among the works it cites.
Natural language object retrieval
R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell · 2016
Later among the works it cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. Yuille, and K. Murphy · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Skip-thought vectors
R. Kiros, Y. Zhu, R. R. Salakhutdinov, R. Zemel, R. Urtasun, A. Torralba, and S. Fidler · 2015
Cited alongside, same era.
Associating neural word embeddings with deep image representations using fisher vectors
B. Klein, G. Lev, G. Sadeh, and L. Wolf · 2015
Cited alongside, same era.
Unsupervised tube extraction using transductive learning and dense trajectories
M. Marian Puscas, E. Sangineto, D. Culibrk, and N. Sebe · 2015
Cited alongside, same era.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Cited alongside, same era.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Cited alongside, same era.
Later among the works it cites.
Object instance search in videos via spatio-temporal trajectory discovery
J. Meng, J. Yuan, J. Yang, G. Wang, and Y.-P. Tan · 2016
Later among the works it cites.
Modeling context between objects for referring expression
V. K. Nagaraja, V. I. Morariu, and L. S. Davis · 2016
Later among the works it cites.
Multi-region two-stream r-cnn for action detection
X. Peng and C. Schmid · 2016
Later among the works it cites.
Grounding of textual phrases in images by reconstruction
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele · 2016
Later among the works it cites.
Deep learning for detecting multiple space-time action tubes in videos
S. Saha, G. Singh, M. Sapienza, P. H. Torr, and F. Cuzzolin · 2016
Later among the works it cites.
Learning deep structure-preserving image-text embeddings
L. Wang, Y. Li, and S. Lazebnik · 2016
Later among the works it cites.
Structured matching for phrase localization
M. Wang, M. Azab, N. Kojima, R. Mihalcea, and J. Deng · 2016
Later among the works it cites.
Track and segment: An iterative unsupervised approach for video object proposals
F. Xiao and Y. J. Lee · 2016
Later among the works it cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Later among the works it cites.