Fetching the paper…
Reading the bibliography…
Visual feature pyramid has shown its superiority in both effectiveness and efficiency in a wide range of applications.
P. Dollár, R. Appel, S. Belongie, and P. Perona, “Fast feature pyramids for object detection,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 36, no. 8, pp. 1532–1545, 2014
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in European Conference on Computer Vision (ECCV) , 2014
2014
Earlier work this paper cites.
R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2014
2014
Earlier work this paper cites.
R. Girshick, “Fast r-cnn,” in International Conference on Computer Vision (ICCV) , 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Spatial pyramid pooling in deep convolutional networks for visual recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 37, no. 9, pp. 1904–1916, 2015
2015
Earlier work this paper cites.
M. Treml, J. Arjona-Medina, T. Unterthiner, R. Durgesh, F. Friedmann, P. Schuberth, A. Mayr, M. Heusel, M. Hofmarcher, M. Widrich et al. , “Speeding up semantic segmentation for autonomous driving,” in Neural Information Processing Systems (NeurIPS) , 2016
2016
Earlier work this paper cites.
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg, “Ssd: Single shot multibox detector,” in European Conference on Computer Vision (ECCV) , 2016
2016
Earlier work this paper cites.
J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “You only look once: Unified, real-time object detection,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016
2016
Earlier work this paper cites.
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba, “Learning deep features for discriminative localization,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016
2016
Earlier work this paper cites.
J. Dai, Y. Li, K. He, and J. Sun, “R-fcn: Object detection via region-based fully convolutional networks,” in Neural Information Processing Systems (NeurIPS) , 2016
2016
Earlier work this paper cites.
G. Larsson, M. Maire, and G. Shakhnarovich, “Fractalnet: Ultra-deep neural networks without residuals,” arXiv , 2016
2016
Earlier work this paper cites.
S. R. Kaiming He, Xiangyu Zhang and J. Sun, “Deep residual learning for image recognition,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016
2016
Earlier work this paper cites.
M. Havaei, A. Davy, D. Warde-Farley, A. Biard, A. Courville, Y. Bengio, C. Pal, P.-M. Jodoin, and H. Larochelle, “Brain tumor segmentation with deep neural networks,” Medical Image Analysis , vol. 35, pp. 18–31, 2017
2017
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 39, no. 6, pp. 1137–1149, 2017
2017
Earlier work this paper cites.
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Neural Information Processing Systems (NeurIPS) , 2017
2017
Earlier work this paper cites.
Z. Luo, A. Mishra, A. Achkar, J. Eichel, S. Li, and P.-M. Jodoin, “Non-local deep features for salient object detection,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2017
2017
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” Neural Information Processing Systems (NeurIPS) , vol. 25, no. 6, pp. 80–90, 2017
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in International Conference on Computer Vision (ICCV) , 2017
2017
Earlier work this paper cites.
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “Mobilenets: Efficient convolutional neural networks for mobile vision applications,” arXiv , 2017
2017
Earlier work this paper cites.
P. Ramachandran, B. Zoph, and Q. V. Le, “Swish: a self-gated activation function,” arXiv , 2017
2017
Earlier work this paper cites.
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He, “Accurate, large minibatch sgd: Training imagenet in 1 hour,” arXiv , 2017
2017
Earlier work this paper cites.
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He, “Aggregated residual transformations for deep neural networks,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2017
2017
Earlier work this paper cites.
J. Redmon and A. Farhadi, “Yolo9000: better, faster, stronger,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2017
2017
Earlier work this paper cites.
C.-Y. Fu, W. Liu, A. Ranga, A. Tyagi, and A. C. Berg, “Dssd: Deconvolutional single shot detector,” arXiv , 2017
2017
Earlier work this paper cites.
H. Zhang, K. Dana, J. Shi, Z. Zhang, X. Wang, A. Tyagi, and A. Agrawal, “Context encoding for semantic segmentation,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018
2018
Cited alongside, same era.
S. Liu, L. Qi, H. Qin, J. Shi, and J. Jia, “Path aggregation network for instance segmentation,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018
2018
Cited alongside, same era.
X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018
2018
Cited alongside, same era.
Y. Chen, Y. Kalantidis, J. Li, S. Yan, and J. Feng, “Aˆ 2-nets: Double attention networks,” in Neural Information Processing Systems (NeurIPS) , 2018
2018
Cited alongside, same era.
J. Redmon and A. Farhadi, “Yolov3: An incremental improvement,” arXiv , 2018
C.-Y. Wang, H.-Y. M. Liao, Y.-H. Wu, P.-Y. Chen, J.-W. Hsieh, and I.-H. Yeh, “Cspnet: A new backbone that can enhance learning capability of cnn,” in IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) , 2020
2020
Later among the works it cites.
J. Cao, H. Cholakkal, R. M. Anwer, F. S. Khan, Y. Pang, and L. Shao, “D2det: Towards high quality object detection and instance segmentation,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2020
2020
Later among the works it cites.
M. Tan, R. Pang, and Q. V. Le, “Efficientdet: Scalable and efficient object detection,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2020
2020
Later among the works it cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in International Conference on Computer Vision (ICCV) , 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, “mixup: Beyond empirical risk minimization,” in International Conference on Learning Representations (ICLR) , 2018
2018
Cited alongside, same era.
Z.-Q. Zhao, P. Zheng, S.-t. Xu, and X. Wu, “Object detection with deep learning: A review,” IEEE Transactions on Neural Networks and Learning Systems , vol. 30, no. 11, pp. 3212–3232, 2019
2019
Cited alongside, same era.
H. Zhang, H. Zhang, C. Wang, and J. Xie, “Co-occurrent features in semantic segmentation,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2019
2019
Cited alongside, same era.
G. Ghiasi, T.-Y. Lin, and Q. V. Le, “Nas-fpn: Learning scalable feature pyramid architecture for object detection,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2019
2019
Cited alongside, same era.
Y. Cao, J. Xu, S. Lin, F. Wei, and H. Hu, “Gcnet: Non-local networks meet squeeze-excitation networks and beyond,” in International Conference on Computer Vision Workshops (ICCVW) , 2019
2019
Cited alongside, same era.
Q. Zhao, T. Sheng, Y. Wang, Z. Tang, Y. Chen, L. Cai, and H. Ling, “M2det: A single-shot object detector based on multi-level feature pyramid network,” in AAAI Conference on Artificial Intelligence (AAAI) , 2019
2019
Cited alongside, same era.
P. Ramachandran, N. Parmar, A. Vaswani, I. Bello, A. Levskaya, and J. Shlens, “Stand-alone self-attention in vision models,” in Neural Information Processing Systems (NeurIPS) , 2019
2019
Cited alongside, same era.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2021
2021
Later among the works it cites.
R. Strudel, R. Garcia, I. Laptev, and C. Schmid, “Segmenter: Transformer for semantic segmentation,” in International Conference on Computer Vision (ICCV) , 2021
2021
Later among the works it cites.
F. Zhu, Y. Zhu, L. Zhang, C. Wu, Y. Fu, and M. Li, “A unified efficient pyramid transformer for semantic segmentation,” in International Conference on Computer Vision Workshops (ICCVW) , 2021
2021
Later among the works it cites.
D. Zhang, H. Zhang, J. Tang, X.-S. Hua, and Q. Sun, “Self-regulation for semantic segmentation,” in International Conference on Computer Vision (ICCV) , 2021
2021
Later among the works it cites.
Z. Ge, S. Liu, F. Wang, Z. Li, and J. Sun, “Yolox: Exceeding yolo series in 2021,” arXiv , 2021
2021
Later among the works it cites.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in International Conference on Machine Learning (ICML) , 2021
2021
Later among the works it cites.
S. Zheng, J. Lu, H. Zhao, X. Zhu, Z. Luo, Y. Wang, Y. Fu, J. Feng, T. Xiang, P. H. Torr et al. , “Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2021
2021
Later among the works it cites.
A. Vaswani, P. Ramachandran, A. Srinivas, N. Parmar, B. Hechtman, and J. Shlens, “Scaling local self-attention for parameter efficient visual backbones,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2021
2021
Later among the works it cites.
I. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Unterthiner, J. Yung, D. Keysers, J. Uszkoreit, M. Lucic et al. , “Mlp-mixer: An all-mlp architecture for vision,” in Neural Information Processing Systems (NeurIPS) , 2021
2021
Later among the works it cites.
H. Liu, Z. Dai, D. R. So, and Q. V. Le, “Pay attention to mlps,” in Neural Information Processing Systems (NeurIPS) , 2021
2021
Later among the works it cites.
W. Yu, M. Luo, P. Zhou, C. Si, Y. Zhou, X. Wang, J. Feng, and S. Yan, “Metaformer is actually what you need for vision,” arXiv , 2021
2021
Later among the works it cites.
Z. Peng, W. Huang, S. Gu, L. Xie, Y. Wang, J. Jiao, and Q. Ye, “Conformer: Local features coupling global representations for visual recognition,” in International Conference on Computer Vision (ICCV) , 2021
2021
Later among the works it cites.
D. Lian, Z. Yu, X. Sun, and S. Gao, “As-mlp: An axial shifted mlp architecture for vision,” arXiv , 2021
2021
Later among the works it cites.
Z. Chen, J. Zhang, and D. Tao, “Recurrent glimpse-based decoder for detection with transformer,” arXiv , 2021
2021
Later among the works it cites.
Y. Fang, B. Liao, X. Wang, J. Fang, J. Qi, R. Wu, J. Niu, and W. Liu, “You only look at one sequence: Rethinking transformer in vision through object detection,” in Neural Information Processing Systems (NeurIPS) , 2021
2021
Later among the works it cites.
H. Song, D. Sun, S. Chun, V. Jampani, D. Han, B. Heo, W. Kim, and M.-H. Yang, “Vidt: An efficient and effective fully transformer-based object detector,” arXiv , 2021
2021
Later among the works it cites.
C.-Y. Wang, A. Bochkovskiy, and H.-Y. M. Liao, “Scaled-yolov4: Scaling cross stage partial network,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2021
2021
Later among the works it cites.
L. Ru, Y. Zhan, B. Yu, and B. Du, “Learning affinity from attention: End-to-end weakly-supervised semantic segmentation with transformers,” arXiv , 2022
2022
Closest in time.
R. Li, Z. Mai, C. Trabelsi, Z. Zhang, J. Jang, and S. Sanner, “Transcam: Transformer attention-based cam refinement for weakly supervised semantic segmentation,” arXiv , 2022
2022
Closest in time.
Q. Hou, Z. Jiang, L. Yuan, M.-M. Cheng, S. Yan, and J. Feng, “Vision permutator: A permutable mlp-like architecture for visual recognition,” arXiv , 2022
2022
Closest in time.