Fetching the paper…
Reading the bibliography…
Although no specific domain knowledge is considered in the design, plain vision transformers have shown excellent performance in visual recognition tasks.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Annotated facial landmarks in the wild: A large-scale, real-world database for facial landmark localization
M. Koestinger, P. Wohlhart, P. M. Roth, and H. Bischof · 2011
Earlier work this paper cites.
Robust face landmark estimation under occlusion
X. P. Burgos-Artizzu, P. Perona, and P. Dollár · 2013
Earlier work this paper cites.
2d human pose estimation: New benchmark and state of the art analysis
M. Andriluka, L. Pishchulin, P. Gehler, and B. Schiele · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Deeppose: Human pose estimation via deep neural networks
A. Toshev and C. Szegedy · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, J. Dean, et al · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Ai challenger: A large-scale dataset for going deeper in image understanding
J. Wu, H. Zheng, B. Zhao, Y. Li, B. Yan, R. Liang, W. Wang, S. Zhou, G. Lin, Y. Fu, et al · 2017
Earlier work this paper cites.
Deep learning using rectified linear units (relu)
A. F. Agarap · 2018
Earlier work this paper cites.
On the convergence of adam and beyond
S. J. Reddi, S. Kale, and S. Kumar · 2018
Earlier work this paper cites.
Simple baselines for human pose estimation and tracking
B. Xiao, H. Wu, and Y. Wei · 2018
Earlier work this paper cites.
Cross-domain adaptation for animal pose estimation
J. Cao, H. Tang, H.-S. Fang, X. Shen, C. Lu, and Y.-W. Tai · 2019
Earlier work this paper cites.
Crowdpose: Efficient crowded scenes pose estimation and a new benchmark
J. Li, C. Wang, H. Zhu, Y. Mao, H.-S. Fang, and C. Lu · 2019
Earlier work this paper cites.
Rethinking on multi-stage networks for human pose estimation
W. Li, Z. Wang, B. Yin, Q. Peng, Y. Du, T. Xiao, G. Yu, H. Lu, Y. Wei, and J. Sun · 2019
Earlier work this paper cites.
Cascade feature aggregation for human pose estimation
Z. Su, M. Ye, G. Zhang, L. Dai, and J. Sheng · 2019
Earlier work this paper cites.
Deep high-resolution representation learning for human pose estimation
K. Sun, B. Xiao, D. Liu, and J. Wang · 2019
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V. Le · 2019
Earlier work this paper cites.
Pose2seg: Detection free human instance segmentation
S.-H. Zhang, R. Li, X. Dong, P. Rosin, Z. Cai, X. Han, D. Yang, H. Huang, and S.-M. Hu · 2019
Cited alongside, same era.
Adversarial semantic data augmentation for human pose estimation
Y. Bin, X. Cao, X. Chen, Y. Ge, Y. Tai, C. Wang, J. Li, F. Huang, C. Gao, and N. Sang · 2020
Cited alongside, same era.
Learning delicate local representations for multi-person pose estimation
Y. Cai, Z. Wang, Z. Luo, B. Yin, A. Du, H. Wang, X. Zhou, E. Zhou, X. Zhang, and J. Sun · 2020
Cited alongside, same era.
End-to-end object detection with transformers
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko · 2020
Cited alongside, same era.
Openmmlab pose estimation toolbox and benchmark
M. Contributors · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao · 2021
Later among the works it cites.
Vitae: Vision transformer advanced by exploring intrinsic inductive bias
Y. Xu, Q. Zhang, J. Zhang, and D. Tao · 2021
Later among the works it cites.
Transpose: Keypoint localization via transformer
S. Yang, Z. Quan, M. Nie, and W. Yang · 2021
Later among the works it cites.
Ap-10k: A benchmark for animal pose estimation in the wild
H. Yu, Y. Xu, J. Zhang, W. Zhao, Z. Guan, and D. Tao · 2021
Later among the works it cites.
Hrformer: High-resolution transformer for dense prediction
Y. Yuan, R. Fu, L. Huang, W. Lin, C. Zhang, X. Chen, and J. Wang · 2021
Later among the works it cites.
Towards high performance human keypoint detection
J. Zhang, Z. Chen, and D. Tao · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The devil is in the details: Delving into unbiased data processing for human pose estimation
J. Huang, Z. Zhu, F. Guo, and G. Huang · 2020
Cited alongside, same era.
Human in events: A large-scale benchmark for human-centric video analysis in complex events
W. Lin, H. Liu, S. Liu, Y. Li, R. Qian, T. Wang, N. Xu, H. Xiong, G.-J. Qi, and N. Sebe · 2020
Cited alongside, same era.
Distribution-aware coordinate representation for human pose estimation
F. Zhang, X. Zhu, H. Dai, M. Ye, and C. Zhu · 2020
Cited alongside, same era.
Empowering things with intelligence: a survey of the progress, challenges, and opportunities in artificial intelligence of things
J. Zhang and D. Tao · 2020
Cited alongside, same era.
Knowledge distillation: A survey
J. Gou, B. Yu, S. J. Maybank, and D. Tao · 2021
Cited alongside, same era.
Multi-instance pose networks: Rethinking top-down pose estimation
R. Khirodkar, V. Chari, A. Agrawal, and A. Tyagi · 2021
Cited alongside, same era.
Later among the works it cites.
Elsa: Enhanced local self-attention for vision transformer
J. Zhou, P. Wang, F. Wang, Q. Liu, H. Li, and R. Jin · 2021
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Closest in time.
BEit: BERT pre-training of image transformers
H. Bao, L. Dong, S. Piao, and F. Wei · 2022
Closest in time.
Bigdetection: A large-scale benchmark for improved object detector pre-training
L. Cai, Z. Zhang, Y. Zhu, L. Zhang, M. Li, and X. Xue · 2022
Closest in time.
Masked autoencoders are scalable vision learners
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick · 2022
Closest in time.
Exploring plain vision transformer backbones for object detection
Y. Li, H. Mao, R. Girshick, and K. He · 2022
Closest in time.
Mvitv2: Improved multiscale vision transformers for classification and detection
Y. Li, C.-Y. Wu, H. Fan, K. Mangalam, B. Xiong, J. Malik, and C. Feichtenhofer · 2022
Closest in time.
Scaled relu matters for training vision transformers
P. Wang, X. Wang, H. Luo, J. Zhou, Z. Zhou, F. Wang, H. Li, and R. Jin · 2022
Closest in time.
Crossformer: A versatile vision transformer hinging on cross-scale attention
W. Wang, L. Yao, L. Chen, B. Lin, D. Cai, X. He, and W. Liu · 2022
Closest in time.
Apt-36k: A large-scale benchmark for animal pose estimation and tracking
Y. Yang, J. Yang, Y. Xu, J. Zhang, L. Lan, and D. Tao · 2022
Closest in time.
Vitaev2: Vision transformer advanced by exploring inductive bias for image recognition and beyond
Q. Zhang, Y. Xu, J. Zhang, and D. Tao · 2022
Closest in time.
Vsa: Learning varied-size window attention in vision transformers
Q. Zhang, Y. Xu, J. Zhang, and D. Tao · 2022
Closest in time.