Fetching the paper…
Reading the bibliography…
Detection Transformer (DETR) and Deformable DETR have been proposed to eliminate the need for many hand-designed components in object detection while demonstrating good performance as previous complex hand-crafted detectors.
H. W. Kuhn, “The hungarian method for the assignment problem,” Naval research logistics quarterly , vol. 2, no. 1-2, pp. 83–97, 1955
1955
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition . IEEE, 2009, pp. 248–255
2009
Earlier work this paper cites.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in Proceedings of the thirteenth international conference on artificial intelligence and statistics , 2010, pp. 249–256
2010
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in European Conference on Computer Vision , 2014
2014
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al. , “Imagenet large scale visual recognition challenge,” International journal of computer vision , vol. 115, no. 3, pp. 211–252, 2015
2015
Earlier work this paper cites.
A. Dosovitskiy, P. Fischer, E. Ilg, P. Hausser, C. Hazirbas, V. Golkov, P. Van Der Smagt, D. Cremers, and T. Brox, “Flownet: Learning optical flow with convolutional networks,” in Proceedings of the IEEE international conference on computer vision , 2015, pp. 2758–2766
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: towards real-time object detection with region proposal networks,” IEEE transactions on pattern analysis and machine intelligence , vol. 39, no. 6, pp. 1137–1149, 2016
2016
Earlier work this paper cites.
J. Dai, Y. Li, K. He, and J. Sun, “R-fcn: Object detection via region-based fully convolutional networks,” in Advances in neural information processing systems , 2016, pp. 379–387
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Shrivastava, A. Gupta, and R. Girshick, “Training region-based object detectors with online hard example mining,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 761–769
2016
Earlier work this paper cites.
Z. Wu, C. Shen, and A. v. d. Hengel, “High-performance semantic segmentation using very deep fully convolutional networks,” arXiv preprint , 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
R. Stewart, M. Andriluka, and A. Y. Ng, “End-to-end people detection in crowded scenes,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 2325–2333
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 2980–2988
2017
Earlier work this paper cites.
K. Kang, H. Li, J. Yan, X. Zeng, B. Yang, T. Xiao, C. Zhang, Z. Wang, R. Wang, X. Wang et al. , “T-cnn: Tubelets with convolutional neural networks for object detection from videos,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 28, no. 10, pp. 2896–2907, 2017
2017
Earlier work this paper cites.
X. Zhu, Y. Xiong, J. Dai, L. Yuan, and Y. Wei, “Deep feature flow for video recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 2349–2358
2017
Earlier work this paper cites.
X. Zhu, Y. Wang, J. Dai, L. Yuan, and Y. Wei, “Flow-guided feature aggregation for video object detection,” in Proceedings of the IEEE International Conference on Computer Vision , 2017, pp. 408–417
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in neural information processing systems , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei, “Deformable convolutional networks,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 764–773
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 2961–2969
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
C. Feichtenhofer, A. Pinz, and A. Zisserman, “Detect to track and track to detect,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 3057–3065
2017
Earlier work this paper cites.
K. Chen, J. Wang, S. Yang, X. Zhang, Y. Xiong, C. C. Loy, and D. Lin, “Optimizing video object detection via a scale-time lattice,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 7814–7823
2018
Earlier work this paper cites.
S. Wang, Y. Zhou, J. Yan, and Z. Deng, “Fully motion-aware network for video object detection,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 542–557
2018
Earlier work this paper cites.
X. Zhu, J. Dai, L. Yuan, and Y. Wei, “Towards high performance video object detection,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 7210–7218
2018
Earlier work this paper cites.
G. Bertasius, L. Torresani, and J. Shi, “Object detection in video with spatiotemporal sampling networks,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 331–346
2018
Earlier work this paper cites.
M. Liu and M. Zhu, “Mobile video object detection with temporally-aware feature maps,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 5686–5695
2018
Cited alongside, same era.
K. Chen, J. Wang, S. Yang, X. Zhang, Y. Xiong, C. C. Loy, and D. Lin, “Optimizing video object detection via a scale-time lattice,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition , pp. 7814–7823, 2018
2018
Cited alongside, same era.
X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 7794–7803
2018
Cited alongside, same era.
H. Hu, J. Gu, Z. Zhang, J. Dai, and Y. Wei, “Relation networks for object detection,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 3588–3597
2018
Cited alongside, same era.
Y. Chen, Y. Cao, H. Hu, and L. Wang, “Memory enhanced global-local aggregation for video object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 10 337–10 346
2020
Later among the works it cites.
2020
Later among the works it cites.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in European Conference on Computer Vision . Springer, 2020, pp. 213–229
2020
Later among the works it cites.
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. Xiao and Y. J. Lee, “Video object detection with an aligned spatial-temporal memory,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 485–501
2018
Cited alongside, same era.
Z. Tian, C. Shen, H. Chen, and T. He, “Fcos: Fully convolutional one-stage object detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 9627–9636
2019
Cited alongside, same era.
H. Belhassen, H. Zhang, V. Fresse, and E.-B. Bourennane, “Improving video object detection by seq-bbox matching.” in VISIGRAPP (5: VISAPP) , 2019, pp. 226–233
2019
Cited alongside, same era.
C. Guo, B. Fan, J. Gu, Q. Zhang, S. Xiang, V. Prinet, and C. Pan, “Progressive sparse local attention for video object detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 3909–3918
2019
Cited alongside, same era.
H. Deng, Y. Hua, T. Song, Z. Zhang, Z. Xue, R. Ma, N. Robertson, and H. Guan, “Object guided external memory network for video object detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 6678–6687
2019
Cited alongside, same era.
Z. Jiang, P. Gao, C. Guo, Q. Zhang, S. Xiang, and C. Pan, “Video object detection with locally-weighted deformable neighbors,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 33, no. 01, 2019, pp. 8529–8536
2019
Cited alongside, same era.
J. Deng, Y. Pan, T. Yao, W. Zhou, H. Li, and T. Mei, “Relation distillation networks for video object detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 7023–7032
2019
Cited alongside, same era.
M. Shvets, W. Liu, and A. C. Berg, “Leveraging long-range temporal relationships between proposals for video object detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 9756–9764
2019
Cited alongside, same era.
2020
Later among the works it cites.
Y. Qian, L. Yu, W. Liu, G. Kang, and A. G. Hauptmann, “Adaptive feature aggregation for video object detection,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision Workshops , 2020, pp. 143–147
2020
Later among the works it cites.
M. Contributors, “MMTracking: OpenMMLab video perception toolbox and benchmark,” https://github.com/open-mmlab/mmtracking , 2020
2020
Later among the works it cites.
F. Wang, Z. Xu, Y. Gan, C.-M. Vong, and Q. Liu, “Scnet: Scale-aware coupling-structure network for efficient video object detection,” Neurocomputing , vol. 404, pp. 283–293, 2020
2020
Later among the works it cites.
Y. Wu, H. Zhang, Y. Li, Y. Yang, and D. Yuan, “Video object detection guided by object blur evaluation,” IEEE Access , vol. 8, pp. 208 554–208 565, 2020
2020
Later among the works it cites.
Z. Xu, E. Hrustic, and D. Vivet, “Centernet heatmap propagation for real-time video object detection,” in European Conference on Computer Vision . Springer, 2020, pp. 220–234
2020
Later among the works it cites.
G. Sun, Y. Hua, G. Hu, and N. Robertson, “Mamba: Multi-level aggregation via memory bank for video object detection,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 35, no. 3, 2021, pp. 2620–2627
2021
Later among the works it cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” The International Conference on Learning Representations (ICLR) , 2021
2021
Later among the works it cites.
L. He, Q. Zhou, X. Li, L. Niu, G. Cheng, X. Li, W. Liu, Y. Tong, L. Ma, and L. Zhang, “End-to-end video object detection with spatial-temporal transformers,” in Proceedings of the 29th ACM International Conference on Multimedia (ACM MM) , 2021, p. 1507–1516
2021
Later among the works it cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” ICCV , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in International Conference on Machine Learning . PMLR, 2021, pp. 10 347–10 357
2021
Later among the works it cites.
P. Sun, R. Zhang, Y. Jiang, T. Kong, C. Xu, W. Zhan, M. Tomizuka, L. Li, Z. Yuan, C. Wang et al. , “Sparse r-cnn: End-to-end object detection with learnable proposals,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 14 454–14 463
2021
Later among the works it cites.
J. Deng, Y. Pan, T. Yao, W. Zhou, H. Li, and T. Mei, “Minet: Meta-learning instance identifiers for video object detection,” IEEE Transactions on Image Processing , vol. 30, pp. 6879–6891, 2021
2021
Later among the works it cites.
T. Gong, K. Chen, X. Wang, Q. Chu, F. Zhu, D. Lin, N. Yu, and H. Feng, “Temporal roi align for video object recognition,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 35, no. 2, 2021, pp. 1442–1450
2021
Later among the works it cites.
Y. Cui, L. Yan, Z. Cao, and D. Liu, “Tf-blender: Temporal feature blender for video object detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , October 2021, pp. 8138–8147
2021
Later among the works it cites.
L. Han, P. Wang, Z. Yin, F. Wang, and H. Li, “Class-aware feature aggregation network for video object detection,” IEEE Transactions on Circuits and Systems for Video Technology , pp. 1–1, 2021
2021
Later among the works it cites.
R. Jin, G. Lin, C. Wen, J. Wang, and F. Liu, “Feature flow: In-network feature flow estimation for video object detection,” Pattern Recognition , vol. 122, p. 108323, 2022
2022
Closest in time.
2022
Closest in time.
X. Li, W. Zhang, J. Pang, K. Chen, G. Cheng, Y. Tong, and C. C. Loy, “Video k-net: A simple, strong, and unified baseline for video segmentation,” in CVPR , 2022
2022
Closest in time.
X. Li, S. Xu, Y. Yang, G. Cheng, Y. Tong, and D. Tao, “Panoptic-partformer: Learning a unified model for panoptic part segmentation,” in Eur. Conf. Comput. Vis. , 2022
2022
Closest in time.
S. Xu, X. Li, J. Wang, G. Cheng, Y. Tong, and D. Tao, “Fashionformer: A simple, effective and unified baseline for human fashion segmentation and recognition,” in Eur. Conf. Comput. Vis. , 2022
2022
Closest in time.