Fetching the paper…
Reading the bibliography…
Driven by the rapid development of deep learning technology, the YOLO series has set a new benchmark for real-time object detectors.
End-to-End Object Detection with Transformers
Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020 · 2005
Earlier work this paper cites.
PP-YOLO: An effective and efficient implementation of object detector
Long, X.; Deng, K.; Wang, G.; Zhang, Y.; Dang, Q.; Gao, Y.; Shen, H.; Ren, J.; Han, S.; Ding, E.; et al. 2020 · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Deformable DETR: Deformable Transformers for End-to-End Object Detection
Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; and Dai, J. 2020 · 2010
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Lin, T.-Y.; Maire, M.; Belongie, S.; Bourdev, L.; Girshick, R.; Hays, J.; Perona, P.; Ramanan, D.; Zitnick, C. L.; and Dollár, P. 2015 · 2015
Earlier work this paper cites.
SSD: Single Shot MultiBox Detector , 21–37
Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.-Y.; and Berg, A. C. 2016 · 2016
Earlier work this paper cites.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Ren, S.; He, K.; Girshick, R.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Language Modeling with Gated Convolutional Networks
Dauphin, Y. N.; Fan, A.; Auli, M.; and Grangier, D. 2017 · 2017
Earlier work this paper cites.
Densely Connected Convolutional Networks
Huang, G.; Liu, Z.; van der Maaten, L.; and Weinberger, K. Q. 2017 · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
Lin, T.-Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; and Belongie, S. 2017 · 2017
Earlier work this paper cites.
Objects365: A large-scale, high-quality dataset for object detection
Shao, S.; Li, Z.; Zhang, T.; Peng, C.; Yu, G.; Zhang, X.; Li, J.; and Sun, J. 2019 · 2019
Earlier work this paper cites.
EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
Tan, M.; and Le, Q. V. 2020 · 2020
Earlier work this paper cites.
Combining Recurrent, Convolutional, and Continuous-time Models with Linear State-Space Layers
Gu, A.; Johnson, I.; Goel, K.; Saab, K.; Dao, T.; Rudra, A.; and Ré, C. 2021 · 2021
Cited alongside, same era.
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 · 2021
Cited alongside, same era.
Mobilevit: light-weight, general-purpose, and mobile-friendly vision transformer
Mehta, S.; and Rastegari, M. 2021 · 2021
Cited alongside, same era.
Convolutional Gated MLP: Combining Convolutions gMLP
Rajagopal, A.; and Nirmala, V. 2021 · 2021
Cited alongside, same era.
Edgevit: Efficient visual modeling for edge computing
Chen, Z.; Zhong, F.; Luo, Q.; Zhang, X.; and Zheng, Y. 2022 · 2022
Rethinking Vision Transformers for MobileNet Size and Speed
Li, Y.; Hu, J.; Wen, Y.; Evangelidis, G.; Salahi, K.; Wang, Y.; Tulyakov, S.; and Ren, J. 2023 · 2023
Later among the works it cites.
TransNeXt: Robust Foveal Visual Perception for Vision Transformers
Shi, D. 2023 · 2023
Later among the works it cites.
Simplified State Space Layers for Sequence Modeling
Smith, J. T. H.; Warrington, A.; and Linderman, S. W. 2023 · 2023
Later among the works it cites.
Repvit: Revisiting mobile cnn from vit perspective
Wang, A.; Chen, H.; Lin, Z.; Pu, H.; and Ding, G. 2023 · 2023
Later among the works it cites.
YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors
Wang, C.-Y.; Bochkovskiy, A.; and Liao, H.-Y. M. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Efficiently Modeling Long Sequences with Structured State Spaces
Gu, A.; Goel, K.; and Ré, C. 2022 · 2022
Cited alongside, same era.
YOLOv6: A single-stage object detection framework for industrial applications
Li, C.; Li, L.; Jiang, H.; Weng, K.; Geng, Y.; Li, L.; Ke, Z.; Li, Q.; Cheng, M.; Nie, W.; et al. 2022 · 2022
Cited alongside, same era.
A ConvNet for the 2020s
Liu, Z.; Mao, H.; Wu, C.-Y.; Feichtenhofer, C.; Darrell, T.; and Xie, S. 2022 · 2022
Cited alongside, same era.
DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection
Zhang, H.; Li, F.; Liu, S.; Zhang, L.; Su, H.; Zhu, J.; Ni, L. M.; and Shum, H.-Y. 2022 · 2022
Cited alongside, same era.
YOLO-MS: Rethinking Multi-Scale Representation Learning for Real-time Object Detection
Chen, Y.; Yuan, X.; Wu, R.; Wang, J.; Hou, Q.; and Cheng, M.-M. 2023 · 2023
Cited alongside, same era.
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Gu, A.; and Dao, T. 2023 · 2023
Cited alongside, same era.
Ultralytics YOLO
Jocher, G.; Chaurasia, A.; and Qiu, J. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
Rethinking mobile block for efficient attention-based models
Zhang, J.; Li, X.; Li, J.; Liu, L.; Xue, Z.; Zhang, B.; Jiang, Z.; Huang, T.; Wang, Y.; and Wang, C. 2023 · 2023
Later among the works it cites.
DETRs Beat YOLOs on Real-time Object Detection
Zhao, Y.; Lv, W.; Xu, S.; Wei, J.; Wang, G.; Dang, Q.; Liu, Y.; and Chen, J. 2023 · 2023
Later among the works it cites.
LocalMamba: Visual State Space Model with Windowed Selective Scan
Huang, T.; Pei, X.; You, S.; Wang, F.; Qian, C.; and Xu, C. 2024 · 2024
Closest in time.
VMamba: Visual State Space Model
Liu, Y.; Tian, Y.; Zhao, Y.; Yu, H.; Xie, L.; Wang, Y.; Ye, Q.; and Liu, Y. 2024 · 2024
Closest in time.
Gold-YOLO: Efficient object detector via gather-and-distribute mechanism
Wang, C.; He, W.; Nie, Y.; Guo, J.; Liu, C.; Wang, Y.; and Han, K. 2024 · 2024
Closest in time.
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
Zhu, L.; Liao, B.; Zhang, Q.; Wang, X.; Liu, W.; and Wang, X. 2024 · 2024
Closest in time.