R. J. Wang, X. Li, and C. X. Ling, “Pelee: A real-time object detection system on mobile devices,” in Advances in Neural Information Processing Systems 31 , S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, Eds. Curran Associates, Inc., 2018, pp. 1967–1976
1976
Earlier work this paper cites.
Y. LeCun, J. S. Denker, and S. A. Solla, “Optimal brain damage,” in Advances in neural information processing systems , 1990, pp. 598–605
1990
Earlier work this paper cites.
X. Zhang, J. Zou, X. Ming, K. He, and J. Sun, “Efficient and accurate approximations of nonlinear convolutional networks,” in CVPR , 2015, pp. 1984–1992
1992
Earlier work this paper cites.
R. Vaillant, C. Monrocq, and Y. Le Cun, “Original approach for the localisation of objects in images,” IEE Proceedings-Vision, Image and Signal Processing , vol. 141, no. 4, pp. 245–250, 1994
1994
Earlier work this paper cites.
H. A. Rowley, S. Baluja, and T. Kanade, “Human face detection in visual scenes,” in Advances in Neural Information Processing Systems , 1996, pp. 875–881
1996
Earlier work this paper cites.
T. G. Dietterich, R. H. Lathrop, and T. Lozano-Pérez, “Solving the multiple instance problem with axis-parallel rectangles,” Artificial intelligence , vol. 89, no. 1-2, pp. 31–71, 1997
1997
Earlier work this paper cites.
C. P. Papageorgiou, M. Oren, and T. Poggio, “A general framework for object detection,” in ICCV . IEEE, 1998, pp. 555–562
1998
Earlier work this paper cites.
D. G. Lowe, “Object recognition from local scale-invariant features,” in ICCV , vol. 2. Ieee, 1999, pp. 1150–1157
1999
Earlier work this paper cites.
P. Simard, L. Bottou, P. Haffner, and Y. LeCun, “Boxlets: a fast convolution algorithm for signal processing and neural networks,” in Advances in Neural Information Processing Systems , 1999, pp. 571–577
1999
Earlier work this paper cites.
C. Papageorgiou and T. Poggio, “A trainable system for object detection,” International journal of computer vision , vol. 38, no. 1, pp. 15–33, 2000
2000
Earlier work this paper cites.
P. Viola and M. Jones, “Rapid object detection using a boosted cascade of simple features,” in CVPR , vol. 1. IEEE, 2001, pp. I–I
2001
Earlier work this paper cites.
A. Torralba and P. Sinha, “Detecting faces in impoverished images,” MASSACHUSETTS INST OF TECH CAMBRIDGE ARTIFICIAL INTELLIGENCE LAB, Tech. Rep., 2001
2001
Earlier work this paper cites.
F. Fleuret and D. Geman, “Coarse-to-fine face detection,” International Journal of computer vision , vol. 41, no. 1-2, pp. 85–107, 2001
2001
Earlier work this paper cites.
S. Belongie, J. Malik, and J. Puzicha, “Shape matching and object recognition using shape contexts,” CALIFORNIA UNIV SAN DIEGO LA JOLLA DEPT OF COMPUTER SCIENCE AND ENGINEERING, Tech. Rep., 2002
2002
Earlier work this paper cites.
S. Andrews, I. Tsochantaridis, and T. Hofmann, “Support vector machines for multiple-instance learning,” in Advances in neural information processing systems , 2003, pp. 577–584
2003
Earlier work this paper cites.
P. Viola and M. J. Jones, “Robust real-time face detection,” International journal of computer vision , vol. 57, no. 2, pp. 137–154, 2004
2004
Earlier work this paper cites.
——, “Distinctive image features from scale-invariant keypoints,” International journal of computer vision , vol. 60, no. 2, pp. 91–110, 2004
2004
Earlier work this paper cites.
N. Dalal and B. Triggs, “Histograms of oriented gradients for human detection,” in CVPR , vol. 1. IEEE, 2005, pp. 886–893
2005
Earlier work this paper cites.
F. Porikli, “Integral histogram: A fast way to extract histograms in cartesian spaces,” in CVPR , vol. 1. IEEE, 2005, pp. 829–836
2005
Earlier work this paper cites.
Q. Zhu, M.-C. Yeh, K.-T. Cheng, and S. Avidan, “Fast human detection using a cascade of histograms of oriented gradients,” in CVPR , vol. 2. IEEE, 2006, pp. 1491–1498
2006
Earlier work this paper cites.
P. Felzenszwalb, D. McAllester, and D. Ramanan, “A discriminatively trained, multiscale, deformable part model,” in CVPR . IEEE, 2008, pp. 1–8
2008
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in CVPR . Ieee, 2009, pp. 248–255
2009
Earlier work this paper cites.
P. Dollár, C. Wojek, B. Schiele, and P. Perona, “Pedestrian detection: A benchmark,” in CVPR . IEEE, 2009, pp. 304–311
2009
Earlier work this paper cites.
S. K. Divvala, D. Hoiem, J. H. Hays, A. A. Efros, and M. Hebert, “An empirical study of context in object detection,” in CVPR . IEEE, 2009, pp. 1271–1278
2009
Earlier work this paper cites.
X. Wang, T. X. Han, and S. Yan, “An hog-lbp human detector with partial occlusion handling,” in ICCV . IEEE, 2009, pp. 32–39
2009
Earlier work this paper cites.
P. Dollár, Z. Tu, P. Perona, and S. Belongie, “Integral channel features,” 2009
2009
Earlier work this paper cites.
P. F. Felzenszwalb, R. B. Girshick, and D. McAllester, “Cascade object detection with deformable part models,” in CVPR . IEEE, 2010, pp. 2241–2248
2010
Earlier work this paper cites.
P. F. Felzenszwalb, R. B. Girshick, D. McAllester, and D. Ramanan, “Object detection with discriminatively trained part-based models,” IEEE transactions on pattern analysis and machine intelligence , vol. 32, no. 9, pp. 1627–1645, 2010
2010
Earlier work this paper cites.
M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman, “The pascal visual object classes (voc) challenge,” International journal of computer vision , vol. 88, no. 2, pp. 303–338, 2010
2010
Earlier work this paper cites.
B. Alexe, T. Deselaers, and V. Ferrari, “What is an object?” in CVPR . IEEE, 2010, pp. 73–80
2010
Earlier work this paper cites.
T. Malisiewicz, A. Gupta, and A. A. Efros, “Ensemble of exemplar-svms for object detection and beyond,” in ICCV . IEEE, 2011, pp. 89–96
2011
Earlier work this paper cites.
R. B. Girshick, P. F. Felzenszwalb, and D. A. Mcallester, “Object detection with grammar models,” in Advances in Neural Information Processing Systems , 2011, pp. 442–450
2011
Earlier work this paper cites.
T. Malisiewicz, Exemplar-based representations for object detection, association and beyond . Carnegie Mellon University, 2011
2011
Earlier work this paper cites.
C. Desai, D. Ramanan, and C. C. Fowlkes, “Discriminative models for multi-class object layout,” International journal of computer vision , vol. 95, no. 1, pp. 1–12, 2011
2011
Earlier work this paper cites.
D. Mrowca, M. Rohrbach, J. Hoffman, R. Hu, K. Saenko, and T. Darrell, “Spatial semantic regularisation for large scale object detection,” in ICCV , 2015, pp. 2003–2011
2011
Earlier work this paper cites.
R. B. Girshick, From rigid templates to grammars: Object detection with structured models . Citeseer, 2012
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in neural information processing systems , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
P. Dollar, C. Wojek, B. Schiele, and P. Perona, “Pedestrian detection: An evaluation of the state of the art,” IEEE transactions on pattern analysis and machine intelligence , vol. 34, no. 4, pp. 743–761, 2012
2012
Earlier work this paper cites.
——, “Measuring the objectness of image windows,” IEEE transactions on pattern analysis and machine intelligence , vol. 34, no. 11, pp. 2189–2202, 2012
2012
Earlier work this paper cites.
I. Kokkinos, “Bounding part scores for rapid detection with deformable part models,” in ECCV . Springer, 2012, pp. 41–50
2012
Earlier work this paper cites.
J. R. Uijlings, K. E. Van De Sande, T. Gevers, and A. W. Smeulders, “Selective search for object recognition,” International journal of computer vision , vol. 104, no. 2, pp. 154–171, 2013
2013
Earlier work this paper cites.
P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun, “Overfeat: Integrated recognition, localization and detection using convolutional networks,” arXiv preprint arXiv:1312.6229 , 2013
Original
2013
Earlier work this paper cites.
C. Szegedy, A. Toshev, and D. Erhan, “Deep neural networks for object detection,” in Advances in neural information processing systems , 2013, pp. 2553–2561
2013
Earlier work this paper cites.
M. Mathieu, M. Henaff, and Y. LeCun, “Fast training of convolutional networks through ffts,” arXiv preprint arXiv:1312.5851 , 2013
Original
2013
Earlier work this paper cites.
M. A. Sadeghi and D. Forsyth, “Fast template evaluation with vector quantization,” in Advances in neural information processing systems , 2013, pp. 2949–2957
2013
Earlier work this paper cites.
B. Hariharan, P. Arbeláez, R. Girshick, and J. Malik, “Simultaneous detection and segmentation,” in ECCV . Springer, 2014, pp. 297–312
2014
Earlier work this paper cites.
R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in CVPR , 2014, pp. 580–587
2014
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Spatial pyramid pooling in deep convolutional networks for visual recognition,” in ECCV . Springer, 2014, pp. 346–361
2014
Earlier work this paper cites.
M. A. Sadeghi and D. Forsyth, “30hz object detection with dpm v5,” in ECCV . Springer, 2014, pp. 65–79
2014
Earlier work this paper cites.
M. D. Zeiler and R. Fergus, “Visualizing and understanding convolutional networks,” in ECCV . Springer, 2014, pp. 818–833
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in ECCV . Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
J. Hosang, R. Benenson, and B. Schiele, “How good are detection proposals, really?” arXiv preprint arXiv:1406.6962 , 2014
Original
2014
Earlier work this paper cites.
M.-M. Cheng, Z. Zhang, W.-Y. Lin, and P. Torr, “Bing: Binarized normed gradients for objectness estimation at 300fps,” in CVPR , 2014, pp. 3286–3293
2014
Earlier work this paper cites.
D. Erhan, C. Szegedy, A. Toshev, and D. Anguelov, “Scalable object detection using deep neural networks,” in CVPR , 2014, pp. 2147–2154
2014
Earlier work this paper cites.
R. Rothe, M. Guillaumin, and L. Van Gool, “Non-maximum suppression for object detection by passing messages between windows,” in Asian Conference on Computer Vision . Springer, 2014, pp. 290–306
2014
Earlier work this paper cites.
N. Vasilache, J. Johnson, M. Mathieu, S. Chintala, S. Piantino, and Y. LeCun, “Fast convolutional nets with fbfft: A gpu performance evaluation,” arXiv preprint arXiv:1412.7580 , 2014
Original
2014
Earlier work this paper cites.
G. Cheng, J. Han, P. Zhou, and L. Guo, “Multi-class geospatial object detection and geographic image classification based on collection of part detectors,” ISPRS Journal of Photogrammetry and Remote Sensing , vol. 98, pp. 119–132, 2014
2014
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Advances in neural information processing systems , 2014, pp. 2672–2680
2014
Earlier work this paper cites.
——, “Hypercolumns for object segmentation and fine-grained localization,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 447–456
2015
Earlier work this paper cites.
A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” in CVPR , 2015, pp. 3128–3137
2015
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” nature , vol. 521, no. 7553, p. 436, 2015
2015
Earlier work this paper cites.
R. Girshick, “Fast r-cnn,” in ICCV , 2015, pp. 1440–1448
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” in Advances in neural information processing systems , 2015, pp. 91–99
2015
Earlier work this paper cites.
M. Everingham, S. A. Eslami, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman, “The pascal visual object classes challenge: A retrospective,” International journal of computer vision , vol. 111, no. 1, pp. 98–136, 2015
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al. , “Imagenet large scale visual recognition challenge,” International Journal of Computer Vision , vol. 115, no. 3, pp. 211–252, 2015
2015
Earlier work this paper cites.
S. Gidaris and N. Komodakis, “Object detection via a multi-region and semantic segmentation-aware cnn model,” in ICCV , 2015, pp. 1134–1142
2015
Earlier work this paper cites.
Q. Chen, Z. Song, J. Dong, Z. Huang, Y. Hua, and S. Yan, “Contextualizing object detection and classification,” IEEE transactions on pattern analysis and machine intelligence , vol. 37, no. 1, pp. 13–27, 2015
2015
Earlier work this paper cites.
S. Gupta, B. Hariharan, and J. Malik, “Exploring person context and local scene context for object detection,” arXiv preprint arXiv:1511.08177 , 2015
Original
2015
Earlier work this paper cites.
L. Wan, D. Eigen, and R. Fergus, “End-to-end integration of a convolution network, deformable parts model and non-maximum suppression,” in CVPR , 2015, pp. 851–859
2015
Earlier work this paper cites.
H. Li, Z. Lin, X. Shen, J. Brandt, and G. Hua, “A convolutional neural network cascade for face detection,” in CVPR , 2015, pp. 5325–5334
2015
Earlier work this paper cites.
Z. Cai, M. Saberian, and N. Vasconcelos, “Learning complexity-aware cascades for deep pedestrian detection,” in ICCV , 2015, pp. 3361–3369
2015
Earlier work this paper cites.
S. Han, H. Mao, and W. J. Dally, “Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding,” arXiv preprint arXiv:1510.00149 , 2015
Original
2015
Earlier work this paper cites.
K. He and J. Sun, “Convolutional neural networks at constrained time cost,” in CVPR , 2015, pp. 5353–5360
2015
Earlier work this paper cites.
O. Rippel, J. Snoek, and R. P. Adams, “Spectral representations for convolutional neural networks,” in Advances in neural information processing systems , 2015, pp. 2449–2457
2015
Earlier work this paper cites.
H. Zhu, X. Chen, W. Dai, K. Fu, Q. Ye, and J. Jiao, “Orientation robust object detection in aerial images using deep convolutional neural network,” in ICIP . IEEE, 2015, pp. 3735–3739
2015
Earlier work this paper cites.
J. Gu, Z. Wang, J. Kuen, L. Ma, A. Shahroudy, B. Shuai, T. Liu, X. Wang, L. Wang, G. Wang et al. , “Recent advances in convolutional neural networks,” arXiv preprint arXiv:1512.07108 , 2015
Original
2015
Earlier work this paper cites.
A. Radford, L. Metz, and S. Chintala, “Unsupervised representation learning with deep convolutional generative adversarial networks,” arXiv preprint arXiv:1511.06434 , 2015
Original
2015
Earlier work this paper cites.
J. Dai, K. He, and J. Sun, “Instance-aware semantic segmentation via multi-task network cascades,” in CVPR , 2016, pp. 3150–3158
2016
Earlier work this paper cites.
J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “You only look once: Unified, real-time object detection,” in CVPR , 2016, pp. 779–788
2016
Earlier work this paper cites.