Fetching the paper…
Reading the bibliography…
Vision transformers have achieved great successes in many computer vision tasks.
Clustered pose and nonlinear appearance models for human pose estimation
Sam Johnson and Mark Everingham · 2010
Earlier work this paper cites.
Learning effective human pose estimation from inaccurate annotation
Sam Johnson and Mark Everingham · 2011
Earlier work this paper cites.
Robust face landmark estimation under occlusion
Xavier P Burgos-Artizzu, Pietro Perona, and Piotr Dollár · 2013
Earlier work this paper cites.
Human3. 6m: Large scale datasets and predictive methods for 3d human sensing in natural environments
Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu · 2013
Earlier work this paper cites.
Supervised descent method and its applications to face alignment
Xuehan Xiong and Fernando De la Torre · 2013
Earlier work this paper cites.
2d human pose estimation: New benchmark and state of the art analysis
Mykhaylo Andriluka, Leonid Pishchulin, Peter Gehler, and Bernt Schiele · 2014
Earlier work this paper cites.
Face alignment by explicit shape regression
Xudong Cao, Yichen Wei, Fang Wen, and Jian Sun · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Multi-source deep learning for human pose estimation
Wanli Ouyang, Xiao Chu, and Xiaogang Wang · 2014
Earlier work this paper cites.
Deeppose: Human pose estimation via deep neural networks
Alexander Toshev and Christian Szegedy · 2014
Earlier work this paper cites.
Spectral embedding based facial expression recognition with multiple features
Kaimin Yu, Zhiyong Wang, Markus Hagenbuchner, and David Dagan Feng · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Earlier work this paper cites.
Face alignment by coarse-to-fine shape searching
Shizhan Zhu, Cheng Li, Chen Change Loy, and Xiaoou Tang · 2015
Earlier work this paper cites.
Study on density peaks clustering based on k-nearest neighbors and principal component analysis
Mingjing Du, Shifei Ding, and Hongjie Jia · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Stacked hourglass networks for human pose estimation
Alejandro Newell, Kaiyu Yang, and Jia Deng · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
Realtime multi-person 2d pose estimation using part affinity fields
Zhe Cao, Tomas Simon, Shih-En Wei, and Yaser Sheikh · 2017
Earlier work this paper cites.
Multi-context attention for human pose estimation
Xiao Chu, Wei Yang, Wanli Ouyang, Cheng Ma, Alan L Yuille, and Xiaogang Wang · 2017
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Monocular 3d human pose estimation in the wild using improved cnn supervision
Dushyant Mehta, Helge Rhodin, Dan Casas, Pascal Fua, Oleksandr Sotnychenko, Weipeng Xu, and Christian Theobalt · 2017
Earlier work this paper cites.
Associative embedding: End-to-end learning for joint detection and grouping
Alejandro Newell, Zhiao Huang, and Jia Deng · 2017
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Charles R Qi, Li Yi, Hao Su, and Leonidas J Guibas · 2017
Earlier work this paper cites.
Leveraging intra and inter-dataset variations for robust face alignment
Wenyan Wu and Shuo Yang · 2017
Earlier work this paper cites.
Learning feature pyramids for human pose estimation
Wei Yang, Shuang Li, Wanli Ouyang, Hongsheng Li, and Xiaogang Wang · 2017
Earlier work this paper cites.
Openpose: realtime multi-person 2d pose estimation using part affinity fields
Zhe Cao, Gines Hidalgo, Tomas Simon, Shih-En Wei, and Yaser Sheikh · 2018
Earlier work this paper cites.
Wing loss for robust facial landmark localisation with convolutional neural networks
Zhen-Hua Feng, Josef Kittler, Muhammad Awais, Patrik Huber, and Xiao-Jun Wu · 2018
Cited alongside, same era.
End-to-end recovery of human shape and pose
Angjoo Kanazawa, Michael J Black, David W Jacobs, and Jitendra Malik · 2018
Cited alongside, same era.
A cascaded inception of inception network with attention modulated feature fusion for human pose estimation
Wentao Liu, Jie Chen, Cheng Li, Chen Qian, Xiao Chu, and Xiaolin Hu · 2018
Cited alongside, same era.
Bodynet: Volumetric inference of 3d human body shapes
Gul Varol, Duygu Ceylan, Bryan Russell, Jimei Yang, Ersin Yumer, Ivan Laptev, and Cordelia Schmid · 2018
Cited alongside, same era.
Recovering accurate 3d human pose in the wild using imus and a moving camera
Timo von Marcard, Roberto Henschel, Michael J Black, Bodo Rosenhahn, and Gerard Pons-Moll · 2018
Cited alongside, same era.
3d human mesh regression with dense correspondence
Wang Zeng, Wanli Ouyang, Ping Luo, Wentao Liu, and Xiaogang Wang · 2020
Later among the works it cites.
Random erasing data augmentation
Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2021
Later among the works it cites.
Learning to regress bodies from images using differentiable semantic rendering
Sai Kumar Dwivedi, Nikos Athanasiou, Muhammed Kocabas, and Michael J Black · 2021
Later among the works it cites.
Handsformer: Keypoint transformer for monocular 3d pose estimation ofhands and object in interaction
Shreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, and Vincent Lepetit · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wayne Wu, Chen Qian, Shuo Yang, Quan Wang, Yici Cai, and Qiang Zhou · 2018
Cited alongside, same era.
Simple baselines for human pose estimation and tracking
Bin Xiao, Haiping Wu, and Yichen Wei · 2018
Cited alongside, same era.
Hierarchical graph representation learning with differentiable pooling
Rex Ying, Jiaxuan You, Christopher Morris, Xiang Ren, William L Hamilton, and Jure Leskovec · 2018
Cited alongside, same era.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz · 2018
Cited alongside, same era.
Trb: a novel triplet representation for understanding 2d human body
Haodong Duan, Kwan-Yee Lin, Sheng Jin, Wentao Liu, Chen Qian, and Wanli Ouyang · 2019
Cited alongside, same era.
Single-network whole-body pose estimation
Gines Hidalgo, Yaadhav Raaj, Haroon Idrees, Donglai Xiang, Hanbyul Joo, Tomas Simon, and Yaser Sheikh · 2019
Cited alongside, same era.
Graph sequence recurrent neural network for vision-based freezing of gait detection
Kun Hu, Zhiyong Wang, Wei Wang, Kaylena A Ehgoetz Martens, Liang Wang, Tieniu Tan, Simon JG Lewis, and David Dagan Feng · 2019
Cited alongside, same era.
Exemplar fine-tuning for 3d human model fitting towards in-the-wild 3d human pose estimation
Hanbyul Joo, Natalia Neverova, and Andrea Vedaldi · 2021
Later among the works it cites.
Human pose regression with residual log-likelihood estimation
Jiefeng Li, Siyuan Bian, Ailing Zeng, Can Wang, Bo Pang, Wentao Liu, and Cewu Lu · 2021
Later among the works it cites.
Localvit: Bringing locality to vision transformers
Yawei Li, Kai Zhang, Jiezhang Cao, Radu Timofte, and Luc Van Gool · 2021
Later among the works it cites.
Tokenpose: Learning keypoint tokens for human pose estimation
Yanjie Li, Shoukui Zhang, Zhicheng Wang, Sen Yang, Wankou Yang, Shu-Tao Xia, and Erjin Zhou · 2021
Later among the works it cites.
End-to-end human pose and mesh reconstruction with transformers
Kevin Lin, Lijuan Wang, and Zicheng Liu · 2021
Later among the works it cites.
Mesh graphormer
Kevin Lin, Lijuan Wang, and Zicheng Liu · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
Tfpose: Direct human pose estimation with transformers
Weian Mao, Yongtao Ge, Chunhua Shen, Zhi Tian, Xinlong Wang, and Zhibin Wang · 2021
Later among the works it cites.
Dynamicvit: Efficient vision transformers with dynamic token sparsification
Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Later among the works it cites.
When human pose estimation meets robustness: Adversarial algorithms and benchmarks
Jiahang Wang, Sheng Jin, Wentao Liu, Weizhong Liu, Chen Qian, and Ping Luo · 2021
Later among the works it cites.
Pnp-detr: Towards efficient visual analysis with transformers
Tao Wang, Li Yuan, Yunpeng Chen, Jiashi Feng, and Shuicheng Yan · 2021
Later among the works it cites.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao · 2021
Later among the works it cites.
Not all images are worth 16x16 words: Dynamic vision transformers with adaptive sequence length
Yulin Wang, Rui Huang, Shiji Song, Zeyi Huang, and Gao Huang · 2021
Later among the works it cites.
Cvt: Introducing convolutions to vision transformers
Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang · 2021
Later among the works it cites.
Graph-based 3d multi-person pose estimation using multi-view images
Size Wu, Sheng Jin, Wentao Liu, Lei Bai, Chen Qian, Dong Liu, and Wanli Ouyang · 2021
Later among the works it cites.
Vipnas: Efficient video pose estimation via neural architecture search
Lumin Xu, Yingda Guan, Sheng Jin, Wentao Liu, Chen Qian, Ping Luo, Wanli Ouyang, and Xiaogang Wang · 2021
Later among the works it cites.
Transpose: Towards explainable human pose estimation by transformer
Sen Yang, Zhibin Quan, Mu Nie, and Wankou Yang · 2021
Later among the works it cites.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zihang Jiang, Francis EH Tay, Jiashi Feng, and Shuicheng Yan · 2021
Later among the works it cites.
Hrformer: High-resolution transformer for dense prediction
Yuhui Yuan, Rao Fu, Lang Huang, Weihong Lin, Chao Zhang, Xilin Chen, and Jingdong Wang · 2021
Later among the works it cites.
Vision transformer with progressive sampling
Xiaoyu Yue, Shuyang Sun, Zhanghui Kuang, Meng Wei, Philip HS Torr, Wayne Zhang, and Dahua Lin · 2021
Later among the works it cites.
3d human pose estimation with spatial and temporal transformers
Ce Zheng, Sijie Zhu, Matias Mendieta, Taojiannan Yang, Chen Chen, and Zhengming Ding · 2021
Later among the works it cites.