Fetching the paper…
Reading the bibliography…
Convolutional neural networks (CNNs) have dominated the field of computer vision for nearly a decade due to their strong ability to learn local features.
H. Idrees, I. Saleemi, C. Seibert, and M. Shah, “Multi-source multi-scale counting in extremely dense crowd images,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2013, pp. 2547–2554
2013
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al. , “Imagenet large scale visual recognition challenge,” International journal of computer vision , vol. 115, no. 3, pp. 211–252, 2015
2015
Earlier work this paper cites.
M. Borstel, M. Kandemir, P. Schmidt, M. Rao, K. Rajamani, and F. Hamprecht, “Gaussian process density counting from weak supervision,” in 2016 European Conference on Computer Vision (ECCV) , 2016, pp. 365–380
2016
Earlier work this paper cites.
Y. Zhang, D. Zhou, S. Chen, S. Gao, and Y. Ma, “Single-image crowd counting via multi-column convolutional neural network,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 589–597
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
H. Idrees, M. Tayyab, K. Athrey, D. Zhang, S. Al-Maadeed, N. Rajpoot, and M. Shah, “Composition loss for counting, density map estimation and localization in dense crowds,” in European Conference on Computer Vision (ECCV) , 2018, pp. 532–546
2018
Earlier work this paper cites.
Z. Shi, L. Zhang, Y. Liu, X. Cao, Y. Ye, M.-M. Cheng, and G. Zheng, “Crowd counting with deep negative correlation learning,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2018, pp. 5382–5390
2018
Earlier work this paper cites.
D. Sam, N. Sajjan, H. Maurya, and R. Babu, “Almost unsupervised learning for dense crowd counting,” Proceedings of the AAAI Conference on Artificial Intelligence , vol. 33, pp. 8868–8875, 2019
2019
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations, ICLR , 2019
2019
Cited alongside, same era.
J. Wan, W. Luo, B. Wu, A. B. Chan, and W. Liu, “Residual regression with semantic prior for crowd counting,” in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 4031–4040
2019
Cited alongside, same era.
Z. Ma, X. Wei, X. Hong, and Y. Gong, “Bayesian loss for crowd count estimation with point supervision,” in 2019 IEEE/CVF International Conference on Computer Vision (ICCV) , 2019, pp. 6141–6150
2019
Cited alongside, same era.
Y. Lei, Y. Liu, P. Zhang, and L. Liu, “Towards using count-level weak supervision for crowd counting,” Pattern Recognition , vol. 109, p. 107616, 2020
2020
Cited alongside, same era.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” in IEEE/CVF International Conference on Computer Vision (CVPR) , 2021, pp. 568–578
2021
Later among the works it cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” International Conference on Computer Vision (ICCV) , 2021
2021
Later among the works it cites.
G. Sun, Y. Liu, T. Probst, D. Pani Paudel, N. Popovic, and L. Van Gool, “Boosting crowd counting with transformers,” arXiv e-prints , pp. arXiv–2105, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Yang, G. Li, Z. Wu, L. Su, and N. Sebe, “Weakly-supervised crowd counting learns from sorting rather than locations,” in 2020 European Conference on Computer Vision (ECCV) , 11 2020, pp. 1–17
2020
Cited alongside, same era.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations , 2021
2021
Cited alongside, same era.
X. Chu, Z. Tian, Y. Wang, B. Zhang, H. Ren, X. Wei, H. Xia, and C. Shen, “Twins: Revisiting the design of spatial attention in vision transformers,” in NeurIPS , 2021
2021
Cited alongside, same era.
2021
Later among the works it cites.
2021
Later among the works it cites.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “PVTv2: Improved baselines with pyramid vision transformer,” Computational Visual Media , vol. 8, no. 3, pp. 1–10, 2022
2022
Closest in time.