Fetching the paper…
Reading the bibliography…
Object detection and object tracking are usually treated as two separate processes.
N. Dalal and B. Triggs, “Histograms of oriented gradients for human detection,” in Computer Vision and Pattern Recognition, 2005. CVPR 2005. IEEE Computer Society Conference on , vol. 1. IEEE, 2005, pp. 886–893
2005
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on . IEEE, 2009, pp. 248–255
2009
Earlier work this paper cites.
P. F. Felzenszwalb, R. B. Girshick, D. McAllester, and D. Ramanan, “Object detection with discriminatively trained part-based models,” IEEE transactions on pattern analysis and machine intelligence , vol. 32, no. 9, pp. 1627–1645, 2010
2010
Earlier work this paper cites.
M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman, “The pascal visual object classes (voc) challenge,” International journal of computer vision , vol. 88, no. 2, pp. 303–338, 2010
2010
Earlier work this paper cites.
J. R. Uijlings, K. E. Van De Sande, T. Gevers, and A. W. Smeulders, “Selective search for object recognition,” International journal of computer vision , vol. 104, no. 2, pp. 154–171, 2013
2013
Earlier work this paper cites.
R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2014, pp. 580–587
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
C. L. Zitnick and P. Dollár, “Edge boxes: Locating object proposals from edges,” in European conference on computer vision . Springer, 2014, pp. 391–405
2014
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Spatial pyramid pooling in deep convolutional networks for visual recognition,” in European Conference on Computer Vision . Springer, 2014, pp. 346–361
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in European Conference on Computer Vision . Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
J.-P. Jodoin, G.-A. Bilodeau, and N. Saunier, “Urban tracker: Multiple object tracking in urban mixed traffic,” in Applications of Computer Vision (WACV), 2014 IEEE Winter Conference on . IEEE, 2014, pp. 885–892
2014
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” in Advances in neural information processing systems , 2015, pp. 91–99
2015
Earlier work this paper cites.
R. Girshick, “Fast r-cnn,” in Proceedings of the IEEE International Conference on Computer Vision , 2015, pp. 1440–1448
2015
Earlier work this paper cites.
C. Ma, J.-B. Huang, X. Yang, and M.-H. Yang, “Hierarchical convolutional features for visual tracking,” in Proceedings of the IEEE international conference on computer vision , 2015, pp. 3074–3082
2015
Cited alongside, same era.
2015
Cited alongside, same era.
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg, “Ssd: Single shot multibox detector,” in European conference on computer vision . Springer, 2016, pp. 21–37
2016
Cited alongside, same era.
J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “You only look once: Unified, real-time object detection,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 779–788
2016
Cited alongside, same era.
J. Li, Y. Wei, X. Liang, J. Dong, T. Xu, J. Feng, and S. Yan, “Attentive contexts for object detection,” IEEE Transactions on Multimedia , vol. 19, no. 5, pp. 944–954, 2017
2017
Later among the works it cites.
C. Feichtenhofer, A. Pinz, and A. Zisserman, “Detect to track and track to detect,” in IEEE International Conference on Computer Vision (ICCV) , 2017
2017
Later among the works it cites.
Y. Lu, C. Lu, and C.-K. Tang, “Online video object detection using association lstm,” in IEEE International Conference on Computer Vision (ICCV) , Oct 2017
2017
Later among the works it cites.
K. Kang, H. Li, T. Xiao, W. Ouyang, J. Yan, X. Liu, and X. Wang, “Object detection in videos with tubelet proposal networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2017
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
K. Kang, W. Ouyang, H. Li, and X. Wang, “Object detection from video tubelets with convolutional neural networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 817–825
2016
Cited alongside, same era.
L. Bertinetto, J. Valmadre, J. F. Henriques, A. Vedaldi, and P. H. Torr, “Fully-convolutional siamese networks for object tracking,” in European conference on computer vision . Springer, 2016, pp. 850–865
2016
Cited alongside, same era.
D. Held, S. Thrun, and S. Savarese, “Learning to track at 100 fps with deep regression networks,” in European Conference on Computer Vision . Springer, 2016, pp. 749–765
2016
Cited alongside, same era.
C. Li, A. Chiang, G. Dobler, Y. Wang, K. Xie, K. Ozbay, M. Ghandehari, J. Zhou, and D. Wang, “Robust vehicle tracking for urban traffic videos at intersections,” in Advanced Video and Signal Based Surveillance (AVSS), 2016 13th IEEE International Conference on . IEEE, 2016, pp. 207–213
2016
Cited alongside, same era.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in Computer Vision (ICCV), 2017 IEEE International Conference on . IEEE, 2017, pp. 2980–2988
2017
Cited alongside, same era.
J. Redmon and A. Farhadi, “Yolo9000: Better, faster, stronger,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017
2017
Cited alongside, same era.
S. Tang, Y. Li, L. Deng, and Y. Zhang, “Object localization based on proposal fusion,” IEEE Transactions on Multimedia , vol. 19, no. 9, pp. 2105–2116, 2017
2017
Cited alongside, same era.
S. Saha, G. Singh, and F. Cuzzolin, “Amtnet: Action-micro-tube regression by end-to-end trainable deep architecture,” in IEEE International Conference on Computer Vision (ICCV) , 2017
2017
Later among the works it cites.
R. Hou, C. Chen, and M. Shah, “Tube convolutional neural network (t-cnn) for action detection in videos,” in IEEE International Conference on Computer Vision (ICCV) , 2017
2017
Later among the works it cites.
V. Kalogeiton, P. Weinzaepfel, V. Ferrari, and C. Schmid, “Action Tubelet Detector for Spatio-Temporal Action Localization,” in IEEE International Conference on Computer Vision (ICCV) , 2017
2017
Later among the works it cites.
W. Luo, B. Yang, and R. Urtasun, “Fast and furious: Real time end-to-end 3d detection, tracking and motion forecasting with a single convolutional net,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 3569–3577
2018
Later among the works it cites.
Z. Shao, W. Wu, Z. Wang, W. Du, and C. Li, “Seaships: A large-scale precisely annotated dataset for ship detection,” IEEE Transactions on Multimedia , vol. 20, no. 10, pp. 2593–2604, 2018
2018
Later among the works it cites.
J. Li, X. Liang, J. Li, Y. Wei, T. Xu, J. Feng, and S. Yan, “Multistage object detection with group recursive learning,” IEEE Transactions on Multimedia , vol. 20, no. 7, pp. 1645–1655, 2018
2018
Later among the works it cites.
K. Chen and W. Tao, “Learning linear regression via single convolutional layer for visual object tracking,” IEEE Transactions on Multimedia , 2018
2018
Later among the works it cites.
M. Jaderberg, K. Simonyan, A. Zisserman et al. , “Spatial transformer networks,” in Advances in Neural Information Processing Systems , 2015, pp. 2017–2025
2025
Closest in time.