Fetching the paper…
Reading the bibliography…
Transformer architectures have become the model of choice in natural language processing and are now being introduced into computer vision tasks such as image classification, object detection, and semantic segmentation.
HumanEva: Synchronized video and motion capture dataset and baseline algorithm for evaluation of articulated human motion
L. Sigal, A. Balan, and M. J. Black · 2010
Earlier work this paper cites.
Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments
C. Ionescu, D. Papava, V. Olaru, and C. Sminchisescu · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Deep networks with stochastic depth
Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger · 2016
Earlier work this paper cites.
Realtime multi-person 2d pose estimation using part affinity fields
Zhe Cao, Tomas Simon, Shih-En Wei, and Yaser Sheikh · 2017
Earlier work this paper cites.
Rmpe: Regional multi-person pose estimation
Hao-Shu Fang, Shuqin Xie, Yu-Wing Tai, and Cewu Lu · 2017
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Thomas N. Kipf and Max Welling · 2017
Earlier work this paper cites.
A simple yet effective baseline for 3d human pose estimation
Julieta Martinez, Rayat Hossain, Javier Romero, and James J. Little · 2017
Earlier work this paper cites.
Monocular 3d human pose estimation in the wild using improved cnn supervision
Dushyant Mehta, Helge Rhodin, Dan Casas, Pascal Fua, Oleksandr Sotnychenko, Weipeng Xu, and Christian Theobalt · 2017
Earlier work this paper cites.
Vnect: Real-time 3d human pose estimation with a single rgb camera
Dushyant Mehta, Srinath Sridhar, Oleksandr Sotnychenko, Helge Rhodin, Mohammad Shafiei, Hans-Peter Seidel, Weipeng Xu, Dan Casas, and Christian Theobalt · 2017
Earlier work this paper cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Towards 3d human pose estimation in the wild: A weakly-supervised approach
Xingyi Zhou, Qixing Huang, Xiao Sun, Xiangyang Xue, and Yichen Wei · 2017
Earlier work this paper cites.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli · 2018
Earlier work this paper cites.
Cascaded pyramid network for multi-person pose estimation
Yilun Chen, Zhicheng Wang, Yuxiang Peng, Zhiqiang Zhang, Gang Yu, and Jian Sun · 2018
Earlier work this paper cites.
Cascaded pyramid network for multi-person pose estimation
Yilun Chen, Zhicheng Wang, Yuxiang Peng, Zhiqiang Zhang, Gang Yu, and Jian Sun · 2018
Earlier work this paper cites.
Learning 3d human pose from structure and motion
Rishabh Dabral, Anurag Mundhada, Uday Kusupati, Safeer Afaque, Abhishek Sharma, and Arjun Jain · 2018
Earlier work this paper cites.
Learning 3d human pose from structure and motion
Rishabh Dabral, Anurag Mundhada, Uday Kusupati, Safeer Afaque, Abhishek Sharma, and Arjun Jain · 2018
Cited alongside, same era.
Ordinal depth supervision for 3D human pose estimation
Georgios Pavlakos, Xiaowei Zhou, and Kostas Daniilidis · 2018
Cited alongside, same era.
Exploiting temporal information for 3d human pose estimation
Mir Rayat Imtiaz Hossain and James J. Little · 2018
Cited alongside, same era.
Exploiting spatial-temporal relationships for 3d pose estimation via graph convolutional networks
Y. Cai, L. Ge, J. Liu, J. Cai, T. Cham, J. Yuan, and N. M. Thalmann · 2019
Cited alongside, same era.
Occlusion-aware networks for 3d human pose estimation in video
Y. Cheng, B. Yang, B. Wang, Y. Wending, and R. Tan · 2019
Cited alongside, same era.
Optimizing network structure for 3d human pose estimation
H. Ci, C. Wang, X. Ma, and Y. Wang · 2019
End-to-end human pose and mesh reconstruction with transformers
Kevin Lin, Lijuan Wang, and Zicheng Liu · 2020
Later among the works it cites.
A comprehensive study of weight sharing in graph networks for 3d human pose estimation
Kenkun Liu, Rongqi Ding, Zhiming Zou, Le Wang, and Wei Tang · 2020
Later among the works it cites.
Attention mechanism exploits temporal contexts: Real-time 3d human pose reconstruction
Ruixu Liu, Ju Shen, He Wang, Chen Chen, Sen-ching Cheung, and Vijayan Asari · 2020
Later among the works it cites.
I2l-meshnet: Image-to-lixel prediction network for accurate 3d human pose and mesh estimation from a single rgb image
Gyeongsik Moon and Kyoung Mu Lee · 2020
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On boosting single-frame 3d human pose estimation via monocular videos
Z. Li, X. Wang, F. Wang, and P. Jiang · 2019
Cited alongside, same era.
Trajectory space factorization for deep video-based 3d human pose estimation
Jiahao Lin and Gim Hee Lee · 2019
Cited alongside, same era.
3d human pose estimation in video with temporal convolutions and semi-supervised training
Dario Pavllo, Christoph Feichtenhofer, David Grangier, and Michael Auli · 2019
Cited alongside, same era.
Deep high-resolution representation learning for human pose estimation
Ke Sun, Bin Xiao, Dong Liu, and Jingdong Wang · 2019
Cited alongside, same era.
Learning deep transformer models for machine translation
Qiang Wang, Bei Li, Tong Xiao, Jingbo Zhu, Changliang Li, Derek F Wong, and Lidia S Chao · 2019
Cited alongside, same era.
Cross-modal self-attention network for referring image segmentation
Linwei Ye, Mrigank Rochan, Zhi Liu, and Yang Wang · 2019
Cited alongside, same era.
Later among the works it cites.
Motion guided 3d pose estimation from videos
Jingbo Wang, Sijie Yan, Yuanjun Xiong, and Dahua Lin · 2020
Later among the works it cites.
Transpose: Towards explainable human pose estimation by transformer
Sen Yang, Zhibin Quan, Mu Nie, and Wankou Yang · 2020
Later among the works it cites.
Srnet: Improving generalization in 3d human pose estimation with a split-and-recombine approach
Ailing Zeng, Xiao Sun, Fuyang Huang, Minhao Liu, Qiang Xu, and Stephen Lin · 2020
Later among the works it cites.
Object-occluded human shape and pose estimation from a single color image
Tianshu Zhang, Buzhen Huang, and Yangang Wang · 2020
Later among the works it cites.
Deep learning-based human pose estimation: A survey, 2020
Ce Zheng, Wenhan Wu, Taojiannan Yang, Sijie Zhu, Chen Chen, Ruixu Liu, Ju Shen, Nasser Kehtarnavaz, and Mubarak Shah · 2020
Later among the works it cites.
Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers
Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip HS Torr, et al · 2020
Later among the works it cites.
Towards deeper graph neural networks with differentiable group normalization
Kaixiong Zhou, Xiao Huang, Yuening Li, Daochen Zha, Rui Chen, and Xia Hu · 2020
Later among the works it cites.
Deformable detr: Deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai · 2020
Later among the works it cites.
Anatomy-aware 3d human pose estimation with bone-based pose decomposition
Tianlang Chen, Chen Fang, Xiaohui Shen, Yiheng Zhu, Zhili Chen, and Jiebo Luo · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Closest in time.
Transformers in vision: A survey
Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, and Mubarak Shah · 2021
Closest in time.