Fetching the paper…
Reading the bibliography…
Human-centric perception tasks, e.g., pedestrian detection, skeleton-based action recognition, and pose estimation, have wide industrial applications, such as metaverse and sports analysis.
Articulated human detection with flexible mixtures of parts
Yi Yang and Deva Ramanan · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Paper doll parsing: Retrieving similar styles to parse clothing items
Kota Yamaguchi, M Hadi Kiapour, and Tamara L Berg · 2013
Earlier work this paper cites.
Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments
Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
2d human pose estimation: New benchmark and state of the art analysis
Mykhaylo Andriluka, Leonid Pishchulin, Peter Gehler, and Bernt Schiele · 2014
Earlier work this paper cites.
Person Attribute Recognition with a Jointly-trained Holistic CNN Model
Patrick Sudowe, Hannah Spitzer, and Bastian Leibe · 2015
Earlier work this paper cites.
Scalable person re-identification: A benchmark
Liang Zheng, Liyue Shen, Lu Tian, Shengjin Wang, Jingdong Wang, and Qi Tian · 2015
Earlier work this paper cites.
Clothes co-parsing via joint image segmentation and labeling with application to clothing retrieval
Xiaodan Liang, Liang Lin, Wei Yang, Ping Luo, Junshi Huang, and Shuicheng Yan · 2016
Earlier work this paper cites.
Ntu rgb+ d: A large scale dataset for 3d human activity analysis
Amir Shahroudy, Jun Liu, Tian-Tsong Ng, and Gang Wang · 2016
Earlier work this paper cites.
Human attribute recognition by deep hierarchical contexts
Yining Li, Chen Huang, Chen Change Loy, and Xiaoou Tang · 2016
Earlier work this paper cites.
Fast open-world person re-identification
Xiatian Zhu, Botong Wu, Dongcheng Huang, and Wei-Shi Zheng · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Look into person: Self-supervised structure-sensitive learning and a new benchmark for human parsing
Ke Gong, Xiaodan Liang, Dongyu Zhang, Xiaohui Shen, and Liang Lin · 2017
Earlier work this paper cites.
Hydraplus-net: Attentive deep features for pedestrian analysis
Xihui Liu, Haiyu Zhao, Maoqing Tian, Lu Sheng, Jing Shao, Shuai Yi, Junjie Yan, and Xiaogang Wang · 2017
Earlier work this paper cites.
Person search with natural language description
Shuang Li, Tong Xiao, Hongsheng Li, Bolei Zhou, Dayu Yue, and Xiaogang Wang · 2017
Earlier work this paper cites.
Monocular 3d human pose estimation in the wild using improved cnn supervision
Dushyant Mehta, Helge Rhodin, Dan Casas, Pascal Fua, Oleksandr Sotnychenko, Weipeng Xu, and Christian Theobalt · 2017
Earlier work this paper cites.
Citypersons: A diverse dataset for pedestrian detection
Shanshan Zhang, Rodrigo Benenson, and Bernt Schiele · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Earlier work this paper cites.
Unite the people: Closing the loop between 3d and 2d human representations
Christoph Lassner, Javier Romero, Martin Kiefel, Federica Bogo, Michael J Black, and Peter V Gehler · 2017
Earlier work this paper cites.
Doublefusion: Real-time capture of human performances with inner body shapes from a single depth sensor
Tao Yu, Zerong Zheng, Kaiwen Guo, Jianhui Zhao, Qionghai Dai, Hao Li, Gerard Pons-Moll, and Yebin Liu · 2018
Earlier work this paper cites.
Neural machine translation with deep attention
Biao Zhang, Deyi Xiong, and Jinsong Su · 2018
Earlier work this paper cites.
Spatial temporal graph convolutional networks for skeleton-based action recognition
Sijie Yan, Yuanjun Xiong, and Dahua Lin · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Recovering accurate 3d human pose in the wild using imus and a moving camera
Timo von Marcard, Roberto Henschel, Michael Black, Bodo Rosenhahn, and Gerard Pons-Moll · 2018
Earlier work this paper cites.
Instance-level human parsing via part grouping network
Ke Gong, Xiaodan Liang, Yicheng Li, Yimin Chen, Ming Yang, and Liang Lin · 2018
Earlier work this paper cites.
Crowdhuman: A benchmark for detecting human in a crowd
Shuai Shao, Zijian Zhao, Boxun Li, Tete Xiao, Gang Yu, Xiangyu Zhang, and Jian Sun · 2018
Earlier work this paper cites.
Posetrack: A benchmark for human pose estimation and tracking
Mykhaylo Andriluka, Umar Iqbal, Eldar Insafutdinov, Leonid Pishchulin, Anton Milan, Juergen Gall, and Bernt Schiele · 2018
Earlier work this paper cites.
The eurocity persons dataset: A novel benchmark for object detection
Markus Braun, Sebastian Krebs, Fabian Flohr, and Dariu M Gavrila · 2018
Earlier work this paper cites.
Resound: Towards action recognition without representation bias
Yingwei Li, Yi Li, and Nuno Vasconcelos · 2018
Earlier work this paper cites.
Single-shot multi-person 3d pose estimation from monocular rgb
Dushyant Mehta, Oleksandr Sotnychenko, Franziska Mueller, Weipeng Xu, Srinath Sridhar, Gerard Pons-Moll, and Christian Theobalt · 2018
Earlier work this paper cites.
Simple baselines for human pose estimation and tracking
Bin Xiao, Haiping Wu, and Yichen Wei · 2018
Earlier work this paper cites.
Modanet: A large-scale street fashion dataset with polygon annotations
Shuai Zheng, Fan Yang, M Hadi Kiapour, and Robinson Piramuthu · 2018
Earlier work this paper cites.
Adaptive temporal encoding network for video instance-level human parsing
Qixian Zhou, Xiaodan Liang, Ke Gong, and Liang Lin · 2018
Earlier work this paper cites.
Independently recurrent neural network (indrnn): Building a longer and deeper rnn
Shuai Li, Wanqing Li, Chris Cook, Ce Zhu, and Yanbo Gao · 2018
Earlier work this paper cites.
Learning clip representations for skeleton-based 3d action recognition
Qiuhong Ke, Mohammed Bennamoun, Senjian An, Ferdous Sohel, and Farid Boussaid · 2018
Earlier work this paper cites.
Chao Li, Qiaoyong Zhong, Di Xie, and Shiliang Pu · 2018
Earlier work this paper cites.
End-to-end recovery of human shape and pose
Angjoo Kanazawa, Michael J Black, David W Jacobs, and Jitendra Malik · 2018
Earlier work this paper cites.
Lcr-net++: Multi-person 2d and 3d pose detection in natural images
Gregory Rogez, Philippe Weinzaepfel, and Cordelia Schmid · 2019
Earlier work this paper cites.
Decomposing the immeasurable sport: A deep learning expected possession value framework for soccer
Javier Fernández, Luke Bornn, and Dan Cervone · 2019
Earlier work this paper cites.
Actions speak louder than goals: Valuing player actions in soccer
Tom Decroos, Lotte Bransen, Jan Van Haaren, and Jesse Davis · 2019
Earlier work this paper cites.
Iiot-based fatigue life indication using augmented reality
Mohamed Khalil, Christoph Bergs, Theodoros Papadopoulos, Roland Wüchner, Kai-Uwe Bletzinger, and Michael Heizmann · 2019
Earlier work this paper cites.
Deep high-resolution representation learning for human pose estimation
Ke Sun, Bin Xiao, Dong Liu, and Jingdong Wang · 2019
Earlier work this paper cites.
Two-stream adaptive graph convolutional networks for skeleton-based action recognition
Lei Shi, Yifan Zhang, Jian Cheng, and Hanqing Lu · 2019
Earlier work this paper cites.
Large-scale datasets for going deeper in image understanding
Jiahong Wu, He Zheng, Bo Zhao, Yixin Li, Baoming Yan, Rui Liang, Wenjia Wang, Shipei Zhou, Guosen Lin, Yanwei Fu, et al · 2019
Earlier work this paper cites.
A richly annotated pedestrian dataset for person retrieval in real surveillance scenarios
Dangwei Li, Zhang Zhang, Xiaotang Chen, and Kaiqi Huang · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Widerperson: A diverse dataset for dense pedestrian detection in the wild
Shifeng Zhang, Yiliang Xie, Jun Wan, Hansheng Xia, Stan Z Li, and Guodong Guo · 2019
Earlier work this paper cites.
Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding
Jun Liu, Amir Shahroudy, Mauricio Perez, Gang Wang, Ling-Yu Duan, and Alex C Kot · 2019
Earlier work this paper cites.
Deepfashion2: A versatile benchmark for detection, pose estimation, segmentation and re-identification of clothing images
Yuying Ge, Ruimao Zhang, Xiaogang Wang, Xiaoou Tang, and Ping Luo · 2019
Earlier work this paper cites.
Aist dance video database: Multi-genre, multi-dancer, and multi-camera database for dance information processing
Shuhei Tsuchida, Satoru Fukayama, Masahiro Hamasaki, and Masataka Goto · 2019
Earlier work this paper cites.
Wider face and pedestrian challenge 2018: Methods and results
Chen Change Loy, Dahua Lin, Wanli Ouyang, Yuanjun Xiong, Shuo Yang, Qingqiu Huang, Dongzhan Zhou, Wei Xia, Quanquan Li, Ping Luo, et al · 2019
Earlier work this paper cites.
Actional-structural graph convolutional networks for skeleton-based action recognition
Maosen Li, Siheng Chen, Xu Chen, Ya Zhang, Yanfeng Wang, and Qi Tian · 2019
Earlier work this paper cites.
Convolutional mesh regression for single-image human shape reconstruction
Nikos Kolotouros, Georgios Pavlakos, and Kostas Daniilidis · 2019
Earlier work this paper cites.
Learning to reconstruct 3d human pose and shape via model-fitting in the loop
Nikos Kolotouros, Georgios Pavlakos, Michael J Black, and Kostas Daniilidis · 2019
Earlier work this paper cites.
Semantic-aware occlusion-robust network for occluded person re-identification
Xiaokang Zhang, Yan Yan, Jing-Hao Xue, Yang Hua, and Hanzi Wang · 2020
Earlier work this paper cites.
Graph neural networks: A review of methods and applications
Jie Zhou, Ganqu Cui, Shengding Hu, Zhengyan Zhang, Cheng Yang, Zhiyuan Liu, Lifeng Wang, Changcheng Li, and Maosong Sun · 2020
Earlier work this paper cites.
Skeleton-based action recognition with shift graph convolutional network
Ke Cheng, Yifan Zhang, Xiangyu He, Weihan Chen, Jian Cheng, and Hanqing Lu · 2020
Earlier work this paper cites.
Decoupled spatial-temporal attention network for skeleton-based action-gesture recognition
Lei Shi, Yifan Zhang, Jian Cheng, and Hanqing Lu · 2020
Earlier work this paper cites.
Self-correction for human parsing
Peike Li, Yunqiu Xu, Yunchao Wei, and Yi Yang · 2020
Earlier work this paper cites.
Finegym: A hierarchical video dataset for fine-grained action understanding
Dian Shao, Yue Zhao, Bo Dai, and Dahua Lin · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
A study of checkpointing in large scale training of deep neural networks
Elvis Rojas, Albert Njoroge Kahira, Esteban Meneses, Leonardo Bautista Gomez, and Rosa M Badia · 2020
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Cited alongside, same era.
Detection in crowded scenes: One proposal, multiple predictions
Xuangeng Chu, Anlin Zheng, Xiangyu Zhang, and Jian Sun · 2020
Cited alongside, same era.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Later among the works it cites.
Exploring plain vision transformer backbones for object detection
Yanghao Li, Hanzi Mao, Ross Girshick, and Kaiming He · 2022
Later among the works it cites.
Vitpose: Simple vision transformer baselines for human pose estimation
Yufei Xu, Jing Zhang, Qiming Zhang, and Dacheng Tao · 2022
Later among the works it cites.
Learning disentangled attribute representations for robust pedestrian attribute recognition
Jian Jia, Naiyu Gao, Fei He, Xiaotang Chen, and Kaiqi Huang · 2022
Later among the works it cites.
Revisiting skeleton-based action recognition
Haodong Duan, Yue Zhao, Kai Chen, Dahua Lin, and Bo Dai · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Matthieu Lin, Chuming Li, Xingyuan Bu, Ming Sun, Chen Lin, Junjie Yan, Wanli Ouyang, and Zhidong Deng · 2020
Cited alongside, same era.
Deformable detr: Deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai · 2020
Cited alongside, same era.
Openmmlab pose estimation toolbox and benchmark
MMPose Contributors · 2020
Cited alongside, same era.
Disentangling and unifying graph convolutions for skeleton-based action recognition
Ziyu Liu, Hongwen Zhang, Zhenghao Chen, Zhiyong Wang, and Wanli Ouyang · 2020
Cited alongside, same era.
Whole-body human pose estimation in the wild
Sheng Jin, Lumin Xu, Jin Xu, Can Wang, Wentao Liu, Chen Qian, Wanli Ouyang, and Ping Luo · 2020
Cited alongside, same era.
Learning semantic neural tree for human parsing
Ruyi Ji, Dawei Du, Libo Zhang, Longyin Wen, Yanjun Wu, Chen Zhao, Feiyue Huang, and Siwei Lyu · 2020
Cited alongside, same era.
Part-aware context network for human parsing
Xiaomei Zhang, Yingying Chen, Bingke Zhu, Jinqiao Wang, and Ming Tang · 2020
Cited alongside, same era.
General multi-label image classification with transformers
Jack Lanchantin, Tianlu Wang, Vicente Ordonez, and Yanjun Qi · 2020
Cited alongside, same era.
Zhihao Li, Jianzhuang Liu, Zhensong Zhang, Songcen Xu, and Youliang Yan · 2022
Later among the works it cites.
Learning visibility for robust dense human body estimation
Chun-Han Yao, Jimei Yang, Duygu Ceylan, Yi Zhou, Yang Zhou, and Ming-Hsuan Yang · 2022
Later among the works it cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Later among the works it cites.
mplug: Effective and efficient vision-language learning by cross-modal skip-connections
Chenliang Li, Haiyang Xu, Junfeng Tian, Wei Wang, Ming Yan, Bin Bi, Jiabo Ye, Hehong Chen, Guohai Xu, Zheng Cao, et al · 2022
Later among the works it cites.
Anchor detr: Query design for transformer-based detector
Yingming Wang, Xiangyu Zhang, Tong Yang, and Jian Sun · 2022
Later among the works it cites.
Masked autoencoders as spatiotemporal learners
Christoph Feichtenhofer, Yanghao Li, Kaiming He, et al · 2022
Later among the works it cites.
Not all tokens are equal: Human-centric visual analysis via token clustering transformer
Wang Zeng, Sheng Jin, Wentao Liu, Chen Qian, Ping Luo, Wanli Ouyang, and Xiaogang Wang · 2022
Later among the works it cites.
Occluded human mesh recovery
Rawal Khirodkar, Shashank Tripathi, and Kris Kitani · 2022
Later among the works it cites.
Learning to estimate robust 3d human mesh from in-the-wild crowded scenes
Hongsuk Choi, Gyeongsik Moon, JoonKyu Park, and Kyoung Mu Lee · 2022
Later among the works it cites.
Motionbert: A unified perspective on learning human motion representations
Wentao Zhu, Xiaoxuan Ma, Zhaoyang Liu, Libin Liu, Wayne Wu, and Yizhou Wang · 2023
Closest in time.
Unihcp: A unified model for human-centric perceptions
Yuanzheng Ci, Yizhou Wang, Meilin Chen, Shixiang Tang, Lei Bai, Feng Zhu, Rui Zhao, Fengwei Yu, Donglian Qi, and Wanli Ouyang · 2023
Closest in time.
Smpl: A skinned multi-person linear model
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black · 2023
Closest in time.
Improving video surveillance systems in banks using deep learning techniques
Mohammad Zahrawi and Khaled Shaalan · 2023
Closest in time.
Crowdq: Predicting the queue state of hospital emergency department using crowdsensing mobility data-driven models
Tieqi Shou, Zhuohan Ye, Yayao Hong, Zhiyuan Wang, Hang Zhu, Zhihan Jiang, Dingqi Yang, Binbin Zhou, Cheng Wang, and Longbiao Chen · 2023
Closest in time.
Understanding user behavior in volumetric video watching: Dataset, analysis and prediction
Kaiyuan Hu, Haowen Yang, Yili Jin, Junhua Liu, Yongting Chen, Miao Zhang, and Fangxin Wang · 2023
Closest in time.
Shape-aware text-driven layered video editing
Yao-Chih Lee, Ji-Ze Genevieve Jang, Yi-Ting Chen, Elizabeth Qiu, and Jia-Bin Huang · 2023
Closest in time.
Diffusion video autoencoders: Toward temporally consistent face video editing via disentangled video encoding
Gyeongman Kim, Hajin Shim, Hyunsu Kim, Yunjey Choi, Junho Kim, and Eunho Yang · 2023
Closest in time.
Beyond appearance: a semantic controllable self-supervised learning framework for human-centric visual tasks
Weihua Chen, Xianzhe Xu, Jian Jia, Hao Luo, Yaohua Wang, Fan Wang, Rong Jin, and Xiuyu Sun · 2023
Closest in time.
Humanbench: Towards general human-centric perception with projector assisted pretraining
Shixiang Tang, Cheng Chen, Qingsong Xie, Meilin Chen, Yizhou Wang, Yuanzheng Ci, Lei Bai, Feng Zhu, Haiyang Yang, Li Yi, et al · 2023
Closest in time.
Hap: Structure-aware masked image modeling for human-centric perception
Junkun Yuan, Xinyu Zhang, Hao Zhou, Jian Wang, Zhongwei Qiu, Zhiyin Shao, Shaofeng Zhang, Sifan Long, Kun Kuang, Kun Yao, et al · 2023
Closest in time.
Smpler-x: Scaling up expressive human pose and shape estimation
Zhongang Cai, Wanqi Yin, Ailing Zeng, Chen Wei, Qingping Sun, Yanjun Wang, Hui En Pang, Haiyi Mei, Mingyuan Zhang, Lei Zhang, et al · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Internlm: A multilingual language model with progressively enhanced capabilities, 2023
InternLM Team · 2023
Closest in time.
Effective whole-body pose estimation with two-stages distillation
Zhendong Yang, Ailing Zeng, Chun Yuan, and Yu Li · 2023
Closest in time.
Yibo Zhou, Hai-Miao Hu, Jinzuo Yu, Zhenbo Xu, Weiqing Lu, and Yuran Cao · 2023
Closest in time.
Explicit box detection unifies end-to-end multi-person pose estimation
Jie Yang, Ailing Zeng, Shilong Liu, Feng Li, Ruimao Zhang, and Lei Zhang · 2023
Closest in time.
Rethinking pose estimation in crowds: overcoming the detection information bottleneck and ambiguity
Mu Zhou, Lucas Stoffl, Mackenzie Weygandt Mathis, and Alexander Mathis · 2023
Closest in time.
Imagebind: One embedding space to bind them all
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra · 2023
Closest in time.
Unipad: A universal pre-training paradigm for autonomous driving
Honghui Yang, Sha Zhang, Di Huang, Xiaoyang Wu, Haoyi Zhu, Tong He, Shixiang Tang, Hengshuang Zhao, Qibo Qiu, Binbin Lin, et al · 2023
Closest in time.
Gd-mae: generative decoder for mae pre-training on lidar point clouds
Honghui Yang, Tong He, Jiaheng Liu, Hua Chen, Boxi Wu, Binbin Lin, Xiaofei He, and Wanli Ouyang · 2023
Closest in time.
Smaug: Sparse masked autoencoder for efficient video-language pre-training
Yuanze Lin, Chen Wei, Huiyu Wang, Alan Yuille, and Cihang Xie · 2023
Closest in time.
Images speak in images: A generalist painter for in-context visual learning
Xinlong Wang, Wen Wang, Yue Cao, Chunhua Shen, and Tiejun Huang · 2023
Closest in time.
Seggpt: Segmenting everything in context
Xinlong Wang, Xiaosong Zhang, Yue Cao, Wen Wang, Chunhua Shen, and Tiejun Huang · 2023
Closest in time.
Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark
Zhenfei Yin, Jiong Wang, Jianjian Cao, Zhelun Shi, Dingning Liu, Mukai Li, Lu Sheng, Lei Bai, Xiaoshui Huang, Zhiyong Wang, et al · 2023
Closest in time.
Shikra: Unleashing multimodal llm’s referential dialogue magic
Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, and Rui Zhao · 2023
Closest in time.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Closest in time.
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Closest in time.
Otter: A multi-modal model with in-context instruction tuning
Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Jingkang Yang, and Ziwei Liu · 2023
Closest in time.
What matters in training a gpt4-style language model with multimodal inputs?
Yan Zeng, Hanbo Zhang, Jiani Zheng, Jiangnan Xia, Guoqiang Wei, Yang Wei, Yuchen Zhang, and Tao Kong · 2023
Closest in time.
Pandagpt: One model to instruction-follow them all
Yixuan Su, Tian Lan, Huayang Li, Jialu Xu, Yan Wang, and Deng Cai · 2023
Closest in time.
Pointllm: Empowering large language models to understand point clouds
Runsen Xu, Xiaolong Wang, Tai Wang, Yilun Chen, Jiangmiao Pang, and Dahua Lin · 2023
Closest in time.
Ulip: Learning a unified representation of language, images, and point clouds for 3d understanding
Le Xue, Mingfei Gao, Chen Xing, Roberto Martín-Martín, Jiajun Wu, Caiming Xiong, Ran Xu, Juan Carlos Niebles, and Silvio Savarese · 2023
Closest in time.
Llark: A multimodal foundation model for music
Josh Gardner, Simon Durand, Daniel Stoller, and Rachel M Bittner · 2023
Closest in time.
Motiongpt: Finetuned llms are general-purpose motion generators
Yaqi Zhang, Di Huang, Bin Liu, Shixiang Tang, Yan Lu, Lu Chen, Lei Bai, Qi Chu, Nenghai Yu, and Wanli Ouyang · 2023
Closest in time.
Llama-adapter: Efficient fine-tuning of language models with zero-init attention
Renrui Zhang, Jiaming Han, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, Peng Gao, and Yu Qiao · 2023
Closest in time.
Ponderv2: Pave the way for 3d foundataion model with a universal pre-training paradigm
Haoyi Zhu, Honghui Yang, Xiaoyang Wu, Di Huang, Sha Zhang, Xianglong He, Tong He, Hengshuang Zhao, Chunhua Shen, Yu Qiao, et al · 2023
Closest in time.
Ponder: Point cloud pre-training via neural rendering
Di Huang, Sida Peng, Tong He, Honghui Yang, Xiaowei Zhou, and Wanli Ouyang · 2023
Closest in time.
Plip: Language-image pre-training for person representation learning
Jialong Zuo, Changqian Yu, Nong Sang, and Changxin Gao · 2023
Closest in time.
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al · 2023
Closest in time.
Vitpose++: Vision transformer for generic body pose estimation
Yufei Xu, Jing Zhang, Qiming Zhang, and Dacheng Tao · 2023
Closest in time.
Skeletonmae: graph-based masked autoencoder for skeleton sequence pre-training
Hong Yan, Yang Liu, Yushen Wei, Zhen Li, Guanbin Li, and Liang Lin · 2023
Closest in time.
Rtmpose: Real-time multi-person pose estimation based on mmpose
Tao Jiang, Peng Lu, Li Zhang, Ningsheng Ma, Rui Han, Chengqi Lyu, Yining Li, and Kai Chen · 2023
Closest in time.
Jrdb-pose: A large-scale dataset for multi-person pose estimation and tracking
Edward Vendrow, Duy Tho Le, Jianfei Cai, and Hamid Rezatofighi · 2023
Closest in time.
Sequential modeling enables scalable learning for large vision models
Yutong Bai, Xinyang Geng, Karttikeya Mangalam, Amir Bar, Alan L Yuille, Trevor Darrell, Jitendra Malik, and Alexei A Efros · 2024
Closest in time.
Instructdiffusion: A generalist modeling interface for vision tasks
Zigang Geng, Binxin Yang, Tiankai Hang, Chen Li, Shuyang Gu, Ting Zhang, Jianmin Bao, Zheng Zhang, Houqiang Li, Han Hu, et al · 2024
Closest in time.