Fetching the paper…
Reading the bibliography…
Searching persons in large-scale image databases with the query of natural language description has important applications in video surveillance.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Evaluating appearance models for recognition, reacquisition, and tracking
D. Gray, S. Brennan, and H. Tao · 2007
Earlier work this paper cites.
Automated flower classification over a large number of classes
M.-E. Nilsback and A. Zisserman · 2008
Earlier work this paper cites.
Attribute-based people search in surveillance environments
D. A. Vaquero, R. S. Feris, D. Tran, L. Brown, A. Hampapur, and M. Turk · 2009
Earlier work this paper cites.
Caltech-ucsd birds 200
P. Welinder, S. Branson, T. Mita, C. Wah, F. Schroff, S. Belongie, and P. Perona · 2010
Earlier work this paper cites.
Person re-identification by probabilistic relative distance comparison
W.-S. Zheng, S. Gong, and T. Xiang · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Human reidentification with transferred metric learning
W. Li, R. Zhao, and X. Wang · 2012
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, T. Mikolov, et al · 2013
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
M. Hodosh, P. Young, and J. Hockenmaier · 2013
Earlier work this paper cites.
Learning a recurrent visual representation for image caption generation
X. Chen and C. L. Zitnick · 2014
Earlier work this paper cites.
Pedestrian attribute recognition at far distance
Y. Deng, P. Luo, C. C. Loy, and X. Tang · 2014
Earlier work this paper cites.
Deepreid: Deep filter pairing neural network for person re-identification
W. Li, R. Zhao, T. Xiao, and X. Wang · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Deep captioning with multimodal recurrent neural networks (m-rnn)
J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. Yuille · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier · 2014
Cited alongside, same era.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh · 2015
Cited alongside, same era.
Microsoft coco captions: Data collection and evaluation server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Cited alongside, same era.
From captions to visual concepts and back
H. Fang, S. Gupta, F. Iandola, R. K. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. C. Platt, et al · 2015
Cited alongside, same era.
Are you talking to a machine? dataset and methods for multilingual image question
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu · 2015
Cited alongside, same era.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2015
Later among the works it cites.
Person re-identification meets image search
L. Zheng, L. Shen, L. Tian, S. Wang, J. Bu, and Q. Tian · 2015
Later among the works it cites.
Simple baseline for visual question answering
B. Zhou, Y. Tian, S. Sukhbaatar, A. Szlam, and R. Fergus · 2015
Later among the works it cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Later among the works it cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Johnson, A. Karpathy, and L. Fei-Fei · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Cited alongside, same era.
Person re-identification by local maximal occurrence representation and metric learning
S. Liao, Y. Hu, X. Zhu, and S. Z. Li · 2015
Cited alongside, same era.
Multi-task deep visual-semantic embedding for video thumbnail selection
W. Liu, T. Mei, Y. Zhang, C. Che, and J. Luo · 2015
Cited alongside, same era.
Ask your neurons: A neural-based approach to answering questions about images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Cited alongside, same era.
Image question answering using convolutional neural network with dynamic parameter prediction
H. Noh, P. H. Seo, and B. Han · 2015
Cited alongside, same era.
Exploring models and data for image question answering
M. Ren, R. Kiros, and R. Zemel · 2015
Cited alongside, same era.
Later among the works it cites.
Segmentation from natural language expressions
R. Hu, M. Rohrbach, and T. Darrell · 2016
Later among the works it cites.
Natural language object retrieval
R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell · 2016
Later among the works it cites.
T-cnn: Tubelets with convolutional neural networks for object detection from videos
K. Kang, H. Li, J. Yan, X. Zeng, B. Yang, T. Xiao, C. Zhang, Z. Wang, R. Wang, X. Wang, et al · 2016
Later among the works it cites.
Object detection from video tubelets with convolutional neural networks
K. Kang, W. Ouyang, H. Li, and X. Wang · 2016
Later among the works it cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2016
Later among the works it cites.
Learning deep representations of fine-grained visual descriptions
S. Reed, Z. Akata, B. Schiele, and H. Lee · 2016
Later among the works it cites.
Dualnet: Domain-invariant network for visual question answering
K. Saito, A. Shin, Y. Ushiku, and T. Harada · 2016
Later among the works it cites.
Deep attributes driven multi-camera person re-identification
C. Su, S. Zhang, J. Xing, W. Gao, and Q. Tian · 2016
Later among the works it cites.
End-to-end deep learning for person search
T. Xiao, S. Li, B. Wang, L. Lin, and X. Wang · 2016
Later among the works it cites.
Object detection in videos with tubelet proposal networks
K. Kang, H. Li, T. Xiao, W. Ouyang, J. Yan, X. Liu, and X. Wang · 2017
Closest in time.