Fetching the paper…
Reading the bibliography…
In-context learning provides a new perspective for multi-task modeling for vision and NLP.
Coupled action recognition and pose estimation from multiple views
Angela Yao, Juergen Gall, and Luc Van Gool · 2012
Earlier work this paper cites.
Human3. 6m: Large scale datasets and predictive methods for 3d human sensing in natural environments
Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu · 2013
Earlier work this paper cites.
Joint learning for attribute-consistent person re-identification
Sameh Khamis, Cheng-Hao Kuo, Vivek K Singh, Vinay D Shet, and Larry S Davis · 2015
Earlier work this paper cites.
Pedestrian detection aided by deep learning semantic tasks
Yonglong Tian, Ping Luo, Xiaogang Wang, and Xiaoou Tang · 2015
Earlier work this paper cites.
Stacked hourglass networks for human pose estimation
Alejandro Newell, Kaiyu Yang, and Jia Deng · 2016
Earlier work this paper cites.
Ntu rgb+ d: A large scale dataset for 3d human activity analysis
Amir Shahroudy, Jun Liu, Tian-Tsong Ng, and Gang Wang · 2016
Earlier work this paper cites.
Multi-task learning with low rank attribute embedding for multi-camera person re-identification
Chi Su, Fan Yang, Shiliang Zhang, Qi Tian, Larry Steven Davis, and Wen Gao · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Deep video generation, prediction and completion of human action sequences
Haoye Cai, Chunyan Bai, Yu-Wing Tai, and Chi-Keung Tang · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Human semantic parsing for person re-identification
Mahdi M Kalayeh, Emrah Basaran, Muhittin Gökmen, Mustafa E Kamasak, and Mubarak Shah · 2018
Earlier work this paper cites.
Convolutional sequence to sequence model for human dynamics
Chen Li, Zhen Zhang, Wee Sun Lee, and Gim Hee Lee · 2018
Earlier work this paper cites.
Look into person: Joint body parsing & pose estimation network and a new benchmark
Xiaodan Liang, Ke Gong, Xiaohui Shen, and Liang Lin · 2018
Earlier work this paper cites.
Multi-grained deep feature learning for pedestrian detection
Chunze Lin, Jiwen Lu, and Jie Zhou · 2018
Earlier work this paper cites.
2d/3d pose estimation and action recognition using multitask deep learning
Diogo C Luvizon, David Picard, and Hedi Tabia · 2018
Earlier work this paper cites.
Mutual learning to adapt for joint human parsing and pose estimation
Xuecheng Nie, Jiashi Feng, and Shuicheng Yan · 2018
Earlier work this paper cites.
Recovering accurate 3d human pose in the wild using imus and a moving camera
Timo Von Marcard, Roberto Henschel, Michael J Black, Bodo Rosenhahn, and Gerard Pons-Moll · 2018
Earlier work this paper cites.
Exploiting spatial-temporal relationships for 3d pose estimation via graph convolutional networks
Yujun Cai, Liuhao Ge, Jun Liu, Jianfei Cai, Tat-Jen Cham, Junsong Yuan, and Nadia Magnenat Thalmann · 2019
Earlier work this paper cites.
Optimizing network structure for 3d human pose estimation
Hai Ci, Chunyu Wang, Xiaoxuan Ma, and Yizhou Wang · 2019
Earlier work this paper cites.
Human motion prediction via spatio-temporal inpainting
Alejandro Hernandez, Jurgen Gall, and Francesc Moreno-Noguer · 2019
Earlier work this paper cites.
Amass: Archive of motion capture as surface shapes
Naureen Mahmood, Nima Ghorbani, Nikolaus F Troje, Gerard Pons-Moll, and Michael J Black · 2019
Earlier work this paper cites.
Learning trajectory dependencies for human motion prediction
Wei Mao, Miaomiao Liu, Mathieu Salzmann, and Hongdong Li · 2019
Earlier work this paper cites.
3d human pose estimation in video with temporal convolutions and semi-supervised training
Dario Pavllo, Christoph Feichtenhofer, David Grangier, and Michael Auli · 2019
Cited alongside, same era.
Convolutional sequence generation for skeleton-based action synthesis
Sijie Yan, Zhizhong Li, Yuanjun Xiong, Huahan Yan, and Dahua Lin · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
3d human pose estimation using spatio-temporal networks with explicit occlusion training
Yu Cheng, Bo Yang, Bo Wang, and Robby T Tan · 2020
Cited alongside, same era.
Learning dynamic relationships for 3d human motion prediction
Qiongjie Cui, Huaijiang Sun, and Fei Yang · 2020
Cited alongside, same era.
Robust motion in-betweening
Pretrained diffusion models for unified human motion synthesis
Jianxin Ma, Shuai Bai, and Chang Zhou · 2022
Later among the works it cites.
Masked autoencoders for point cloud self-supervised learning
Yatian Pang, Wenxiao Wang, Francis EH Tay, Wei Liu, Yonghong Tian, and Li Yuan · 2022
Later among the works it cites.
P-stmo: Pre-trained spatial temporal many-to-one model for 3d human pose estimation
Wenkang Shan, Zhenhua Liu, Xinfeng Zhang, Shanshe Wang, Siwei Ma, and Wen Gao · 2022
Later among the works it cites.
Fashionformer: A simple, effective and unified baseline for human fashion segmentation and recognition
Shilin Xu, Xiangtai Li, Jingbo Wang, Guangliang Cheng, Yunhai Tong, and Dacheng Tao · 2022
Later among the works it cites.
Mixste: Seq2seq mixed spatio-temporal encoder for 3d human pose estimation in video
Jinlu Zhang, Zhigang Tu, Jianyu Yang, Yujin Chen, and Junsong Yuan · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Félix G Harvey, Mike Yurick, Derek Nowrouzezahrai, and Christopher Pal · 2020
Cited alongside, same era.
Moglow: Probabilistic and controllable motion synthesis using normalising flows
Gustav Eje Henter, Simon Alexanderson, and Jonas Beskow · 2020
Cited alongside, same era.
Convolutional autoencoders for human motion infilling
Manuel Kaufmann, Emre Aksan, Jie Song, Fabrizio Pece, Remo Ziegler, and Otmar Hilliges · 2020
Cited alongside, same era.
A comprehensive study of weight sharing in graph networks for 3d human pose estimation
Kenkun Liu, Rongqi Ding, Zhiming Zou, Le Wang, and Wei Tang · 2020
Cited alongside, same era.
History repeats itself: Human motion prediction via motion attention
Wei Mao, Miaomiao Liu, and Mathieu Salzmann · 2020
Cited alongside, same era.
Motion guided 3d pose estimation from videos
Jingbo Wang, Sijie Yan, Yuanjun Xiong, and Dahua Lin · 2020
Cited alongside, same era.
Generative tweening: Long-term inbetweening of 3d human motions
Yi Zhou, Jingwan Lu, Connelly Barnes, Jimei Yang, Sitao Xiang, et al · 2020
Cited alongside, same era.
Graformer: Graph-oriented transformer for 3d pose estimation
Weixi Zhao, Weiqiang Wang, and Yunjie Tian · 2022
Later among the works it cites.
Sequential modeling enables scalable learning for large vision models, 2023
Yutong Bai, Xinyang Geng, Karttikeya Mangalam, Amir Bar, Alan Yuille, Trevor Darrell, Jitendra Malik, and Alexei A Efros · 2023
Closest in time.
Towards in-context scene understanding
Ivana Balažević, David Steiner, Nikhil Parthasarathy, Relja Arandjelović, and Olivier J Hénaff · 2023
Closest in time.
Unihcp: A unified model for human-centric perceptions
Yuanzheng Ci, Yizhou Wang, Meilin Chen, Shixiang Tang, Lei Bai, Feng Zhu, Rui Zhao, Fengwei Yu, Donglian Qi, and Wanli Ouyang · 2023
Closest in time.
Explore in-context learning for 3d point cloud understanding
Zhongbin Fang, Xiangtai Li, Xia Li, Joachim M Buhmann, Chen Change Loy, and Mengyuan Liu · 2023
Closest in time.
Unified pose sequence modeling
Lin Geng Foo, Tianjiao Li, Hossein Rahmani, Qiuhong Ke, and Jun Liu · 2023
Closest in time.
Diffpose: Toward more reliable 3d pose estimation
Jia Gong, Lin Geng Foo, Zhipeng Fan, Qiuhong Ke, Hossein Rahmani, and Jun Liu · 2023
Closest in time.
Back to mlp: A simple baseline for human motion prediction
Wen Guo, Yuming Du, Xi Shen, Vincent Lepetit, Xavier Alameda-Pineda, and Francesc Moreno-Noguer · 2023
Closest in time.
Transformer-based visual segmentation: A survey
Xiangtai Li, Henghui Ding, Wenwei Zhang, Haobo Yuan, Guangliang Cheng, Pang Jiangmiao, Kai Chen, Ziwei Liu, and Chen Change Loy · 2023
Closest in time.
Exploring effective factors for improving visual in-context learning
Yanpeng Sun, Qiang Chen, Jian Wang, Jingdong Wang, and Zechao Li · 2023
Closest in time.
3d human pose estimation with spatio-temporal criss-cross attention
Zhenhua Tang, Zhaofan Qiu, Yanbin Hao, Richang Hong, and Ting Yao · 2023
Closest in time.
Eqmotion: Equivariant multi-agent motion prediction with invariant interaction reasoning
Chenxin Xu, Robby T Tan, Yuhong Tan, Siheng Chen, Yu Guang Wang, Xinchao Wang, and Yanfeng Wang · 2023
Closest in time.
Gla-gcn: Global-local adaptive graph convolutional network for 3d human pose estimation from monocular video
Bruce XB Yu, Zhi Zhang, Yongxu Liu, Sheng-hua Zhong, Yan Liu, and Chang Wen Chen · 2023
Closest in time.
What makes good examples for visual in-context learning?
Yuanhan Zhang, Kaiyang Zhou, and Ziwei Liu · 2023
Closest in time.
Motionbert: A unified perspective on learning human motion representations
Wentao Zhu, Xiaoxuan Ma, Zhaoyang Liu, Libin Liu, Wayne Wu, and Yizhou Wang · 2023
Closest in time.
Vg4d: Vision-language model goes 4d video recognition
Zhichao Deng, Xiangtai Li, Xia Li, Yunhai Tong, Shen Zhao, and Mengyuan Liu · 2024
Closest in time.