Fetching the paper…
Reading the bibliography…
Traditional feature encoding scheme (e.g., Fisher vector) with local descriptors (e.g., SIFT) and recent convolutional neural networks (CNNs) are two classes of successful methods for image recognition.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, November 1998
1998
Earlier work this paper cites.
A. Oliva and A. Torralba, “Modeling the shape of the scene: A holistic representation of the spatial envelope,” International Journal of Computer Vision , vol. 42, no. 3, pp. 145–175, 2001
2001
Earlier work this paper cites.
J. Sivic and A. Zisserman, “Video google: A text retrieval approach to object matching in videos,” in ICCV , 2003, pp. 1470–1477
2003
Earlier work this paper cites.
D. G. Lowe, “Distinctive image features from scale-invariant keypoints,” International Journal of Computer Vision , vol. 60, no. 2, pp. 91–110, 2004
2004
Earlier work this paper cites.
G. Csurka, C. Dance, L. Fan, J. Willamowski, and C. Bray, “Visual categorization with bags of keypoints,” in ECCVW , 2004
2004
Earlier work this paper cites.
N. Dalal and B. Triggs, “Histograms of oriented gradients for human detection,” in CVPR , 2005, pp. 886–893
2005
Earlier work this paper cites.
M. Aharon, M. Elad, and A. Bruckstein, “k -svd: An algorithm for designing overcomplete dictionaries for sparse representation,” IEEE Transactions on Signal Processing , vol. 54, no. 11, pp. 4311–4322, Nov 2006
2006
Earlier work this paper cites.
S. Lazebnik, C. Schmid, and J. Ponce, “Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories,” in CVPR , 2006, pp. 2169–2178
2006
Earlier work this paper cites.
H. Bay, A. Ess, T. Tuytelaars, and L. J. V. Gool, “Speeded-up robust features (SURF),” Computer Vision and Image Understanding , vol. 110, no. 3, pp. 346–359, 2008
2008
Earlier work this paper cites.
J. C. van Gemert, J. Geusebroek, C. J. Veenman, and A. W. M. Smeulders, “Kernel codebooks for scene categorization,” in ECCV , 2008, pp. 696–709
2008
Earlier work this paper cites.
A. Vedaldi and B. Fulkerson, “VLFeat: An open and portable library of computer vision algorithms,” http://www.vlfeat.org/ , 2008
2008
Earlier work this paper cites.
J. Yang, K. Yu, Y. Gong, and T. S. Huang, “Linear spatial pyramid matching using sparse coding for image classification,” in CVPR , 2009, pp. 1794–1801
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L. Li, K. Li, and F. Li, “ImageNet: A large-scale hierarchical image database,” in CVPR , 2009, pp. 248–255
2009
Earlier work this paper cites.
A. Quattoni and A. Torralba, “Recognizing indoor scenes,” in CVPR , 2009, pp. 413–420
2009
Earlier work this paper cites.
F. Perronnin, J. Sánchez, and T. Mensink, “Improving the fisher kernel for large-scale image classification,” in ECCV , 2010, pp. 143–156
2010
Earlier work this paper cites.
J. Xiao, J. Hays, K. A. Ehinger, A. Oliva, and A. Torralba, “SUN database: Large-scale scene recognition from abbey to zoo,” in CVPR , 2010, pp. 3485–3492
2010
Earlier work this paper cites.
X. Zhou, K. Yu, T. Zhang, and T. S. Huang, “Image classification using super-vector coding of local image descriptors,” in ECCV , 2010, pp. 141–154
2010
Earlier work this paper cites.
J. Wang, J. Yang, K. Yu, F. Lv, T. S. Huang, and Y. Gong, “Locality-constrained linear coding for image classification,” in CVPR , 2010, pp. 3360–3367
2010
Earlier work this paper cites.
Y. Boureau, F. R. Bach, Y. LeCun, and J. Ponce, “Learning mid-level features for recognition,” in CVPR , 2010, pp. 2559–2566
2010
Earlier work this paper cites.
D. Song and D. Tao, “Biologically inspired feature manifold for scene classification,” IEEE Trans. Image Processing , vol. 19, no. 1, pp. 174–184, 2010
2010
Earlier work this paper cites.
L. Li, H. Su, E. P. Xing, and F. Li, “Object bank: A high-level image representation for scene classification & semantic feature sparsification,” in NIPS , 2010, pp. 1378–1386
2010
Earlier work this paper cites.
J. Wu and J. M. Rehg, “CENTRIST: A visual descriptor for scene categorization,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 33, no. 8, pp. 1489–1501, 2011
2011
Earlier work this paper cites.
M. Pandey and S. Lazebnik, “Scene recognition and weakly supervised object localization with deformable part-based models,” in ICCV , 2011, pp. 1307–1314
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in NIPS , 2012, pp. 1106–1114
2012
Earlier work this paper cites.
H. Jégou, F. Perronnin, M. Douze, J. Sánchez, P. Pérez, and C. Schmid, “Aggregating local image descriptors into compact codes,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 34, no. 9, pp. 1704–1716, 2012
2012
Earlier work this paper cites.
S. Singh, A. Gupta, and A. A. Efros, “Unsupervised discovery of mid-level discriminative patches,” in ECCV , 2012, pp. 73–86
2012
Earlier work this paper cites.
S. N. Parizi, J. G. Oberlin, and P. F. Felzenszwalb, “Reconfigurable models for scene recognition,” in CVPR , 2012, pp. 2775–2782
2012
Cited alongside, same era.
X. Wang, L. Wang, and Y. Qiao, “A comparative study of encoding, pooling and normalization methods for action recognition,” in ACCV , 2012, pp. 572–585
2012
Cited alongside, same era.
S. Singh, A. Gupta, and A. A. Efros, “Unsupervised discovery of mid-level discriminative patches,” in ECCV , 2012, pp. 73–86
2012
Cited alongside, same era.
J. Sánchez, F. Perronnin, T. Mensink, and J. J. Verbeek, “Image classification with the fisher vector: Theory and practice,” International Journal of Computer Vision , vol. 105, no. 3, pp. 222–245, 2013
2013
Cited alongside, same era.
J. Yu, D. Tao, Y. Rui, and J. Cheng, “Pairwise constraints based multiview features fusion for scene classification,” Pattern Recognition , vol. 46, no. 2, pp. 483–496, 2013
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in ICML , 2015, pp. 448–456
2015
Later among the works it cites.
M. Dixit, S. Chen, D. Gao, N. Rasiwasia, and N. Vasconcelos, “Scene classification with semantic fisher vectors,” in CVPR , 2015, pp. 2974–2983
2015
Later among the works it cites.
L. Wang, Y. Qiao, and X. Tang, “Action recognition with trajectory-pooled deep-convolutional descriptors,” in CVPR , 2015, pp. 4305–4314
2015
Later among the works it cites.
2015
Later among the works it cites.
S. Yang and D. Ramanan, “Multi-scale recognition with dag-cnns,” in ICCV , 2015, pp. 1215–1223
2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2013
Cited alongside, same era.
M. Juneja, A. Vedaldi, C. V. Jawahar, and A. Zisserman, “Blocks that shout: Distinctive parts for scene classification,” in CVPR , 2013, pp. 923–930
2013
Cited alongside, same era.
T. Kobayashi, “BFO meets HOG: feature extraction based on histograms of oriented p.d.f. gradients for image classification,” in CVPR , 2013, pp. 747–754
2013
Cited alongside, same era.
C. Doersch, A. Gupta, and A. Efros, “Mid-level visual element discovery as discriminative mode seeking,” in NIPS , 2013, pp. 494–502
2013
Cited alongside, same era.
B. Zhou, À. Lapedriza, J. Xiao, A. Torralba, and A. Oliva, “Learning deep features for scene recognition using places database,” in NIPS , 2014, pp. 487–495
2014
Cited alongside, same era.
J. Yu, Y. Rui, and D. Tao, “Click prediction for web image reranking using multimodal sparse coding,” IEEE Trans. Image Processing , vol. 23, no. 5, pp. 2019–2032, 2014
2014
Cited alongside, same era.
V. Sydorov, M. Sakurada, and C. H. Lampert, “Deep fisher kernels - end to end learning of the fisher kernel GMM parameters,” in CVPR , 2014, pp. 1402–1409
2014
Cited alongside, same era.
X. Peng, L. Wang, Y. Qiao, and Q. Peng, “Boosting VLAD with supervised dictionary learning and high-order statistics,” in ECCV , 2014, pp. 660–674
2014
Cited alongside, same era.
Later among the works it cites.
2015
Later among the works it cites.
2015
Later among the works it cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016, pp. 770–778
2016
Closest in time.
L. Shen, Z. Lin, and Q. Huang, “Relay backpropagation for effective learning of deep convolutional neural networks,” in ECCV , 2016, pp. 467–482
2016
Closest in time.
2016
Closest in time.
J. Yu, X. Yang, F. Gao, and D. Tao, “Deep multimodal distance metric learning using click constraints for image ranking,” IEEE Transactions on Cybernetics , pp. 1–11, 2016
2016
Closest in time.
Z. Xu, S. Huang, Y. Zhang, and D. Tao, “Webly-supervised fine-grained visual categorization via deep domain adaptation,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2016
2016
Closest in time.
Z. Wang, Y. Wang, L. Wang, and Y. Qiao, “Codebook enhancement of VLAD representation for visual recognition,” in ICASSP , 2016, pp. 1258–1262
2016
Closest in time.
2016
Closest in time.
2016
Closest in time.
R. Arandjelovic, P. Gronat, A. Torii, T. Pajdla, and J. Sivic, “NetVLAD: CNN architecture for weakly supervised place recognition,” in CVPR , 2016, pp. 5297–5307
2016
Closest in time.
L. Wang, Y. Qiao, and X. Tang, “Mofap: A multi-level representation for action recognition,” International Journal of Computer Vision , vol. 119, no. 3, pp. 254–271, 2016
2016
Closest in time.
L. Wang, Y. Qiao, X. Tang, and L. V. Gool, “Actionness estimation using hybrid fully convolutional networks,” in CVPR , 2016, pp. 2708–2717
2016
Closest in time.
X. Peng, L. Wang, X. Wang, and Y. Qiao, “Bag of visual words and fusion methods for action recognition: Comprehensive study and good practice,” Computer Vision and Image Understanding , vol. 150, pp. 109–125, 2016
2016
Closest in time.
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. V. Gool, “Temporal segment networks: Towards good practices for deep action recognition,” in ECCV , 2016, pp. 20–36
2016
Closest in time.
Z. Zuo, B. Shuai, G. Wang, X. Liu, X. Wang, B. Wang, and Y. Chen, “Learning contextual dependence with convolutional hierarchical recurrent neural networks,” IEEE Trans. Image Processing , vol. 25, no. 7, pp. 2983–2996, 2016
2016
Closest in time.
M. Cimpoi, S. Maji, I. Kokkinos, and A. Vedaldi, “Deep filter banks for texture recognition, description, and segmentation,” International Journal of Computer Vision , vol. 118, no. 1, pp. 65–94, 2016
2016
Closest in time.
Z. Xu, D. Tao, S. Huang, and Y. Zhang, “Friend or foe: Fine-grained categorization with weak supervision,” IEEE Trans. Image Processing , vol. 26, no. 1, pp. 135–146, 2017
2017
Closest in time.
T. Liu, D. Tao, M. Song, and S. J. Maybank, “Algorithm-dependent generalization bounds for multi-task learning,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 39, no. 2, pp. 227–241, 2017
2017
Closest in time.
S. Guo, W. Huang, L. Wang, and Y. Qiao, “Locally supervised deep hybrid model for scene recognition,” IEEE Trans. Image Processing , vol. 26, no. 2, pp. 808–820, 2017
2017
Closest in time.