Fetching the paper…
Reading the bibliography…
The improvements in recent CNN-based object detection works, from R-CNN [11], Fast/Faster R-CNN [10, 31] to recent Mask R-CNN [14] and RetinaNet [24], mainly come from new network, new framework, or novel loss design.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Maxout networks
I. Goodfellow, D. Warde-Farley, M. Mirza, A. Courville, and Y. Bengio · 2013
Earlier work this paper cites.
Selective search for object recognition
J. R. Uijlings, K. E. Van De Sande, T. Gevers, and A. W. Smeulders · 2013
Earlier work this paper cites.
cudnn: Efficient primitives for deep learning
S. Chetlur, C. Woolley, P. Vandermersch, J. Cohen, J. Tran, B. Catanzaro, and E. Shelhamer · 2014
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Earlier work this paper cites.
Spatial pyramid pooling in deep convolutional networks for visual recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2014
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
A. Krizhevsky · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Overfeat: Integrated recognition, localization and detection using convolutional networks
P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun · 2014
Earlier work this paper cites.
Object detection via a multi-region and semantic segmentation-aware cnn model
S. Gidaris and N. Komodakis · 2015
Earlier work this paper cites.
Fast r-cnn
R. Girshick · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Multi-scale context aggregation by dilated convolutions
F. Yu and V. Koltun · 2015
Earlier work this paper cites.
Inside-outside net: Detecting objects in context with skip pooling and recurrent neural networks
S. Bell, C. Lawrence Zitnick, K. Bala, and R. Girshick · 2016
Cited alongside, same era.
L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille · 2016
Cited alongside, same era.
Training deep nets with sublinear memory cost
T. Chen, B. Xu, C. Zhang, and C. Guestrin · 2016
Cited alongside, same era.
Instance-aware semantic segmentation via multi-task network cascades
J. Dai, K. He, and J. Sun · 2016
Cited alongside, same era.
R-fcn: Object detection via region-based fully convolutional networks
J. Dai, Y. Li, K. He, and J. Sun · 2016
Cited alongside, same era.
Mask r-cnn
K. He, G. Gkioxari, P. Dollar, and R. Girshick · 2017
Closest in time.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
E. Hoffer, I. Hubara, and D. Soudry · 2017
Closest in time.
Squeeze-and-excitation networks
J. Hu, L. Shen, and G. Sun · 2017
Closest in time.
Speed/accuracy trade-offs for modern convolutional object detectors
J. Huang, V. Rathod, C. Sun, M. Zhu, A. Korattikara, A. Fathi, I. Fischer, Z. Wojna, Y. Song, S. Guadarrama, et al · 2017
Closest in time.
Attentive contexts for object detection
J. Li, Y. Wei, X. Liang, J. Dong, T. Xu, J. Feng, and S. Yan · 2017
Closest in time.
Feature pyramid networks for object detection
T.-Y. Lin, P. Dollar, R. Girshick, K. He, B. Hariharan, and S. Belongie · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Ssd: Single shot multibox detector
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg · 2016
Cited alongside, same era.
You only look once: Unified, real-time object detection
J. Redmon, S. Divvala, R. Girshick, and A. Farhadi · 2016
Cited alongside, same era.
Yolo9000: better, faster, stronger
J. Redmon and A. Farhadi · 2016
Cited alongside, same era.
Contextual priming and feedback for faster r-cnn
A. Shrivastava and A. Gupta · 2016
Cited alongside, same era.
Training region-based object detectors with online hard example mining
A. Shrivastava, A. Gupta, and R. Girshick · 2016
Cited alongside, same era.
Deformable convolutional networks
J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei · 2017
Cited alongside, same era.
Closest in time.
Focal loss for dense object detection
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollar · 2017
Closest in time.
What can help pedestrian detection?
J. Mao, T. Xiao, Y. Jiang, and Z. Cao · 2017
Closest in time.
Large kernel matters – improve semantic segmentation by global convolutional network
C. Peng, X. Zhang, G. Yu, G. Luo, and J. Sun · 2017
Closest in time.
Object detection networks on convolutional feature maps
S. Ren, K. He, R. Girshick, X. Zhang, and J. Sun · 2017
Closest in time.
Inception-v4, inception-resnet and the impact of residual connections on learning
C. Szegedy, S. Ioffe, V. Vanhoucke, and A. A. Alemi · 2017
Closest in time.
Aggregated residual transformations for deep neural networks
S. Xie, R. Girshick, P. Dollar, Z. Tu, and K. He · 2017
Closest in time.
Y. You, Z. Zhang, C.-J. Hsieh, J. Demmel, and K. Keutzer · 2017
Closest in time.