Fetching the paper…
Reading the bibliography…
Text-to-image generative models can generate high-quality humans, but realism is lost when generating hands.
Continuous sign language recognition: Towards large vocabulary statistical recognition systems handling multiple signers
Oscar Koller, Jens Forster, and Hermann Ney · 2015
Earlier work this paper cites.
SMPL: A skinned multi-person linear model
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Generating images from captions with attention
Elman Mansimov, Emilio Parisotto, Jimmy Ba, and Ruslan Salakhutdinov · 2016
Earlier work this paper cites.
The kit motion-language dataset
Matthias Plappert, Christian Mandery, and Tamim Asfour · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Embodied hands: Modeling and capturing hands and bodies together
Javier Romero, Dimitrios Tzionas, and Michael J. Black · 2017
Earlier work this paper cites.
Text2action: Generative adversarial synthesis from language to action
Hyemin Ahn, Timothy Ha, Yunho Choi, Hwiyeon Yoo, and Songhwai Oh · 2018
Earlier work this paper cites.
Generating animated videos of human activities from natural language descriptions
Angela S Lin, Lemeng Wu, Rodolfo Corona, Kevin Tai, Qixing Huang, and Raymond J Mooney · 2018
Earlier work this paper cites.
Attngan: Fine-grained text to image generation with attentional generative adversarial networks
Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, and Xiaodong He · 2018
Earlier work this paper cites.
Language2pose: Natural language grounded pose forecasting
Chaitanya Ahuja and Louis-Philippe Morency · 2019
Earlier work this paper cites.
Openpose: Realtime multi-person 2d pose estimation using part affinity fields
Z. Cao, G. Hidalgo Martinez, T. Simon, S. Wei, and Y. A. Sheikh · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Amass: Archive of motion capture as surface shapes
Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll, and Michael J. Black · 2019
Earlier work this paper cites.
Contextual attention for hand detection in the wild
Supreeth Narasimhaswamy, Zhengwei Wei, Yang Wang, Justin Zhang, and Minh Hoai · 2019
Earlier work this paper cites.
Efficient learning on point clouds with basis point sets
Sergey Prokudin, Christoph Lassner, and Javier Romero · 2019
Earlier work this paper cites.
WorkingHands: A hand-tool assembly dataset for image segmentation and activity mining
Roy Shilkrot, Supreeth Narasimhaswamy, Saif Vazir, and Minh Hoai · 2019
Earlier work this paper cites.
Dm-gan: Dynamic memory generative adversarial networks for text-to-image synthesis
Minfeng Zhu, Pingbo Pan, Wei Chen, and Yi Yang · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Ntu rgb+d 120: A large-scale benchmark for 3d human activity understanding
Jun Liu, Amir Shahroudy, Mauricio Perez, Gang Wang, Ling-Yu Duan, and Alex C. Kot · 2020
Earlier work this paper cites.
Detecting hands and recognizing physical contact in the wild
Supreeth Narasimhaswamy, Trung Nguyen, and Minh Hoai · 2020
Earlier work this paper cites.
Mediapipe hands: On-device real-time hand tracking
Fan Zhang, Valentin Bazarevsky, Andrey Vakunov, Andrei Tkachenka, George Sung, Chuo-Ling Chang, and Matthias Grundmann · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol · 2021
Earlier work this paper cites.
Synthesis of compositional animations from textual descriptions
Anindita Ghosh, Noshaba Cheema, Cennet Oguz, Christian Theobalt, and Philipp Slusallek · 2021
Earlier work this paper cites.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans · 2021
Earlier work this paper cites.
Action-conditioned 3d human motion synthesis with transformer vae
Mathis Petrovich, Michael J. Black, and Gül Varol · 2021
Earlier work this paper cites.
Babel: Bodies, action and behavior with english labels
Abhinanda R. Punnakkal, Arjun Chandrasekaran, Nikos Athanasiou, Alejandra Quiros-Ramirez, and Michael J. Black · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Cited alongside, same era.
Monocular, One-stage, Regression of Multiple 3D People
Yu Sun, Qian Bao, Wu Liu, Yili Fu, Black Michael J., and Tao Mei · 2021
Cited alongside, same era.
Improving text-to-image synthesis using contrastive learning
Hui Ye, Xiulong Yang, Martin Takac, Rajshekhar Sunderraman, and Shihao Ji · 2021
Cited alongside, same era.
Cross-modal contrastive learning for text-to-image generation
LAION-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade W Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa R Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev · 2022
Later among the works it cites.
Goal: Generating 4d whole-body motion for hand-object grasping
Omid Taheri, Vasileios Choutas, Michael J. Black, and Dimitrios Tzionas · 2022
Later among the works it cites.
Df-gan: A simple and effective baseline for text-to-image synthesis
Ming Tao, Hao Tang, Fei Wu, Xiao-Yuan Jing, Bing-Kun Bao, and Changsheng Xu · 2022
Later among the works it cites.
Clip-actor: Text-driven recommendation and stylization for animating human meshes
Kim Youwang, Kim Ji-Yeon, and Tae-Hyun Oh · 2022
Later among the works it cites.
Motiondiffuse: Text-driven human motion generation with diffusion model
Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, and Ziwei Liu · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Han Zhang, Jing Yu Koh, Jason Baldridge, Honglak Lee, and Yinfei Yang · 2021
Cited alongside, same era.
Pix2seq: A language modeling framework for object detection
Ting Chen, Saurabh Saxena, Lala Li, David J. Fleet, and Geoffrey Hinton · 2022
Cited alongside, same era.
PoseScript: 3D Human Poses from Natural Language
Delmas, Ginger and Weinzaepfel, Philippe and Lucas, Thomas and Moreno-Noguer, Francesc and Rogez, Grégory · 2022
Cited alongside, same era.
Make-a-scene: Scene-based text-to-image generation with human priors
Oran Gafni, Adam Polyak, Oron Ashual, Shelly Sheynin, Devi Parikh, and Yaniv Taigman · 2022
Cited alongside, same era.
Vector quantized diffusion model for text-to-image synthesis
Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo · 2022
Cited alongside, same era.
Generating diverse and natural 3d human motions from text
Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, and Li Cheng · 2022
Cited alongside, same era.
Cascaded diffusion models for high fidelity image generation
Jonathan Ho, Chitwan Saharia, William Chan, David J. Fleet, Mohammad Norouzi, and Tim Salimans · 2022
Cited alongside, same era.
Later among the works it cites.
Toch: Spatio-temporal object-to-hand correspondence for motion refinement
Keyang Zhou, Bharat Lal Bhatnagar, Jan Eric Lenssen, and Gerard Pons-Moll · 2022
Later among the works it cites.
Firefly, https://www.adobe.com/sensei/generative-ai/firefly.html , 2023
Adobe · 2023
Later among the works it cites.
Disco-Diffusion, https://github.com/alembics/disco-diffusion , 2023
Alembics · 2023
Later among the works it cites.
Mofusion: A framework for denoising-diffusion-based motion synthesis
Rishabh Dabral, Muhammad Hamza Mughal, Vladislav Golyanik, and Christian Theobalt · 2023
Later among the works it cites.
PoseFix: Correcting 3D Human Poses with Natural Language
Delmas, Ginger and Weinzaepfel, Philippe and Moreno-Noguer, Francesc and Rogez, Grégory · 2023
Later among the works it cites.
Arctic: A dataset for dexterous bimanual hand-object manipulation
Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J. Black, and Otmar Hilliges · 2023
Later among the works it cites.
Imos: Intent-driven full-body motion synthesis for human-object interactions
Anindita Ghosh, Rishabh Dabral, Vladislav Golyanik, Christian Theobalt, and Philipp Slusallek · 2023
Later among the works it cites.
A systematic survey of prompt engineering on vision-language foundation models
Jindong Gu, Zhen Han, Shuo Chen, Ahmad Beirami, Bailan He, Gengyuan Zhang, Ruotong Liao, Yao Qin, Volker Tresp, and Philip Torr · 2023
Later among the works it cites.
Affordpose: A large-scale dataset of hand-object interactions with affordance-driven hand pose
Juntao Jian, Xiuping Liu, Manyi Li, Ruizhen Hu, and Jian Liu · 2023
Later among the works it cites.
Flame: Free-form language-based motion synthesis and editing
Jihoon Kim, Jiseob Kim, and Sungjoon Choi · 2023
Later among the works it cites.
Handrefiner: Refining malformed hands in generated images by diffusion-based conditional inpainting, 2023
Wenquan Lu, Yufei Xu, Jing Zhang, Chaoyue Wang, and Dacheng Tao · 2023
Later among the works it cites.
https://www.midjourney.com , 2023
Midjourney · 2023
Later among the works it cites.
Text-to-hand-image generation using pose- and mesh-guided diffusion
Supreeth Narasimhaswamy, Uttaran Bhattacharya, Xiang Chen, Ishita Dasgupta, and Saayan Mitra · 2023
Later among the works it cites.
Dall-E 3, https://openai.com/dall-e-3 , 2023
OpenAI · 2023
Later among the works it cites.
Grip: Generating interaction poses using latent consistency and spatial cues
Omid Taheri, Yi Zhou, Dimitrios Tzionas, Yang Zhou, Duygu Ceylan, Soren Pirk, and Michael J Black · 2023
Later among the works it cites.
Human motion diffusion model
Guy Tevet, Sigal Raab, Brian Gordon, Yoni Shafir, Daniel Cohen-or, and Amit Haim Bermano · 2023
Later among the works it cites.
YOLOv8, https://github.com/ultralytics/ultralytics , 2023
Ultralytics · 2023
Later among the works it cites.
Reco: Region-controlled text-to-image generation
Zhengyuan Yang, Jianfeng Wang, Zhe Gan, Linjie Li, Kevin Lin, Chenfei Wu, Nan Duan, Zicheng Liu, Ce Liu, Michael Zeng, and Lijuan Wang · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala · 2023
Later among the works it cites.
Cvt-slr: Contrastive visual-textual transformation for sign language recognition with variational alignment
Jiangbin Zheng, Yile Wang, Cheng Tan, Siyuan Li, Ge Wang, Jun Xia, Yidong Chen, and Stan Z. Li · 2023
Later among the works it cites.
Hoist-former: Hand-held objects identification, segmentation, and tracking in the wild
Supreeth Narasimhaswamy, Huy Nguyen, Lihan Huang, and Minh Hoai · 2024
Closest in time.