Fetching the paper…
Reading the bibliography…
Deep convolutional neural networks (CNNs) have proven highly effective for visual recognition, where learning a universal representation from activations of convolutional layer plays a fundamental problem.
Exploiting generative models in discriminative classifiers
T. S. Jaakkola, D. Haussler, et al · 1998
Earlier work this paper cites.
VLFeat: An open and portable library of computer vision algorithms
A. Vedaldi and B. Fulkerson · 2008
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
P. Vincent, H. Larochelle, Y. Bengio, and P.-A. Manzagol · 2008
Earlier work this paper cites.
Aggregating local descriptors into a compact image representation
H. Jégou, M. Douze, C. Schmid, and P. Pérez · 2010
Earlier work this paper cites.
Improving the fisher kernel for large-scale image classification
F. Perronnin, J. Sánchez, and T. Mensink · 2010
Earlier work this paper cites.
The Caltech-UCSD Birds-200-2011 Dataset
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie · 2011
Earlier work this paper cites.
Random feature maps for dot product kernels
P. Kar and H. Karnick · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
UCF101: A dataset of 101 human action classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
M. D. Zeiler · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
Image classification with the fisher vector: Theory and practice
J. Sánchez, F. Perronnin, T. Mensink, and J. Verbeek · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Earlier work this paper cites.
Bird species categorization using pose normalized deep convolutional nets
S. Branson, G. Van Horn, S. Belongie, and P. Perona · 2014
Earlier work this paper cites.
Spatial pyramid pooling in deep convolutional networks for visual recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2014
Earlier work this paper cites.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Earlier work this paper cites.
Semi-supervised learning with deep generative models
D. P. Kingma, S. Mohamed, D. J. Rezende, and M. Welling · 2014
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
D. J. Rezende, S. Mohamed, and D. Wierstra · 2014
Cited alongside, same era.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Part-based r-cnns for fine-grained category detection
N. Zhang, J. Donahue, R. Girshick, and T. Darrell · 2014
Cited alongside, same era.
Activitynet: A large-scale video benchmark for human activity understanding
F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. Carlos Niebles · 2015
Cited alongside, same era.
Fast r-cnn
R. Girshick · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Later among the works it cites.
Action recognition with trajectory-pooled deep-convolutional descriptors
L. Wang, Y. Qiao, and X. Tang · 2015
Later among the works it cites.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio · 2015
Later among the works it cites.
A discriminative cnn video representation for event detection
Z. Xu, Y. Yang, and A. G. Hauptmann · 2015
Later among the works it cites.
Beyond short snippets: Deep networks for video classification
J. Yue-Hei Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Later among the works it cites.
Deep filter banks for texture recognition, description, and segmentation
M. Cimpoi, S. Maji, I. Kokkinos, and A. Vedaldi · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Spatial transformer networks
M. Jaderberg, K. Simonyan, A. Zisserman, et al · 2015
Cited alongside, same era.
Fine-grained recognition without part annotations
J. Krause, H. Jin, J. Yang, and L. Fei-Fei · 2015
Cited alongside, same era.
Bilinear cnn models for fine-grained visual recognition
T.-Y. Lin, A. RoyChowdhury, and S. Maji · 2015
Cited alongside, same era.
The treasure beneath convolutional layers: Cross-convolutional-layer pooling for image classification
L. Liu, C. Shen, and A. van den Hengel · 2015
Cited alongside, same era.
Fully convolutional networks for semantic segmentation
J. Long, E. Shelhamer, and T. Darrell · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Cited alongside, same era.
Closest in time.
Convolutional two-stream network fusion for video action recognition
C. Feichtenhofer, A. Pinz, and A. Zisserman · 2016
Closest in time.
Compact bilinear pooling
Y. Gao, O. Beijbom, N. Zhang, and T. Darrell · 2016
Closest in time.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Closest in time.
Composing graphical models with neural networks for structured representations and fast inference
M. J. Johnson, D. Duvenaud, A. B. Wiltschko, S. R. Datta, and R. P. Adams · 2016
Closest in time.
Action recognition using visual attention
S. Sharma, R. Kiros, and R. Salakhutdinov · 2016
Closest in time.
Long-term temporal convolutions for action recognition
G. Varol, I. Laptev, and C. Schmid · 2016
Closest in time.
Temporal segment networks: towards good practices for deep action recognition
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool · 2016
Closest in time.
Picking deep filter responses for fine-grained image recognition
X. Zhang, H. Xiong, W. Zhou, W. Lin, and Q. Tian · 2016
Closest in time.
A key volume mining deep framework for action recognition
W. Zhu, J. Hu, G. Sun, X. Cao, and Y. Qiao · 2016
Closest in time.
Denoising criterion for variational framework
D. J. Im, S. Ahn, R. Memisevic, and Y. Bengio · 2017
Closest in time.