Fetching the paper…
Reading the bibliography…
Diffusion models have experienced a surge of interest as highly expressive yet efficiently trainable probabilistic models.
Comparison of Parametric Representations for Monosyllabic Word Recognition in Continuously Spoken Sentences
Steven Davis and Paul Mermelstein. 1980 · 1980
Earlier work this paper cites.
Mixture Density Networks
Christopher M. Bishop. 1994 · 1994
Earlier work this paper cites.
Practical Parameterization of Rotations Using the Exponential Map
F. Sebastian Grassia. 1998 · 1998
Earlier work this paper cites.
BEAT: The Behavior Expression Animation Toolkit. In Annual Conference on Computer Graphics and Interactive Techniques (SIGGRAPH ’01) . ACM, 477–486
Justine Cassell, Hannes Högni Vilhjálmsson, and Timothy Bickmore. 2001 · 2001
Earlier work this paper cites.
Training Products of Experts by Minimizing Contrastive Divergence
Geoffrey E. Hinton. 2002 · 2002
Earlier work this paper cites.
Max – A Multimodal Assistant in Virtual Reality Construction
Stefan Kopp, Bernhard Jung, Nadine Leßmann, and Ipke Wachsmuth. 2003 · 2003
Earlier work this paper cites.
Information Theory, Inference and Learning Algorithms
David J. C. MacKay. 2003 · 2003
Earlier work this paper cites.
Estimation of Non-Normalized Statistical Models by Score Matching
Aapo Hyvärinen and Peter Dayan. 2005 · 2005
Earlier work this paper cites.
Chroma-Based Statistical Audio Features for Audio Matching. In Proceedings of the IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA ’05) . IEEE, 275–278
Meinard Müller, Frank Kurth, and Michael Clausen. 2005 · 2005
Earlier work this paper cites.
Efficient Content-Based Retrieval of Motion Capture Data
Meinard Müller, Tido Röder, and Michael Clausen. 2005 · 2005
Earlier work this paper cites.
Comparison of Some Listening Test Methods: A Case Study
Etienne Parizet, Nacer Hamzaoui, and Guillaume Sabatié. 2005 · 2005
Earlier work this paper cites.
Nonverbal Behavior Generator for Embodied Conversational Agents. In Proceedings of the International Conference on Intelligent Virtual Agents (IVA ’06) . Springer, 243–255
Jina Lee and Stacy Marsella. 2006 · 2006
Earlier work this paper cites.
Learning to Generate Diverse Dance Motions with Transformer
Jiaman Li, Yihang Yin, Hang Chu, Yi Zhou, Tingwu Wang, Sanja Fidler, and Hao Li. 2020 · 2008
Earlier work this paper cites.
FMDistance: A Fast and Effective Distance Function for Motion Capture Data. In Proceedings of the Annual Conference of the European Association for Computer Graphics – Short Papers (EUROGRAPHICS ’08) , Katerina Mania and Eric Reinhard (Eds.). The Eurographics Association
Kensuke Onuma, Christos Faloutsos, and Jessica K. Hodgins. 2008 · 2008
Earlier work this paper cites.
Gesture Controllers
Sergey Levine, Philipp Krähenbühl, Sebastian Thrun, and Vladlen Koltun. 2010 · 2010
Earlier work this paper cites.
Evaluating the Effect of Gesture and Language on Personality Perception in Conversational Agents. In Proceedings of the International Conference on Intelligent Virtual Agents (IVA ’10) . Springer, 222–235
Michael Neff, Yingying Wang, Rob Abbott, and Marilyn Walker. 2010 · 2010
Earlier work this paper cites.
Example-Based Automatic Music-Driven Conventional Dance Motion Synthesis
Rukun Fan, Songhua Xu, and Weidong Geng. 2011 · 2011
Earlier work this paper cites.
Learn2Dance: Learning Statistical Music-to-Dance Mappings for Choreography Synthesis
Ferda Ofli, Engin Erzin, Yücel Yemez, and A. Murat Tekalp. 2011 · 2011
Earlier work this paper cites.
The Effect of Using Normalized Models in Statistical Speech Synthesis. In Proceedings of the Annual Conference of the International Speech Communication Association (Interspeech ’11) . ISCA, 121–124
Matt Shannon, Heiga Zen, and William Byrne. 2011 · 2011
Earlier work this paper cites.
Generation and Evaluation of Communicative Robot Gesture
Maha Salem, Stefan Kopp, Ipke Wachsmuth, Katharina Rohlfing, and Frank Joublin. 2012 · 2012
Earlier work this paper cites.
When What You Hear Influences When You See: Listening to an Auditory Rhythm Influences the Temporal Allocation of Visual Attention
Jared E. Miller, Laura A. Carlson, and J. Devin McAuley. 2013 · 2013
Earlier work this paper cites.
The Effect of Posture and Dynamics on the Perception of Emotion. In Proceedings of the ACM Symposium on Applied Perception (SAP ’13) . ACM, 91–98
Aline Normoyle, Fannie Liu, Mubbasir Kapadia, Norman I. Badler, and Sophie Jörg. 2013 · 2013
Earlier work this paper cites.
Gesture and Speech in Interaction: An Overview
Petra Wagner, Zofia Malisz, and Stefan Kopp. 2014 · 2013
Earlier work this paper cites.
Predicting Co-Verbal Gestures: A Deep and Temporal Modeling Approach. In Proceedings of the International Conference on Intelligent Virtual Agents . 152–166
Chung-Cheng Chiu, Louis-Philippe Morency, and Stacy Marsella. 2015 · 2015
Earlier work this paper cites.
Recurrent Network Models for Human Dynamics. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV ’15) . IEEE, 4346–4354
Katerina Fragkiadaki, Sergey Levine, Panna Felsen, and Jitendra Malik. 2015 · 2015
Earlier work this paper cites.
Music Content Driven Automated Choreography with Beat-Wise Motion Connectivity Constraints
Satoru Fukayama and Masataka Goto. 2015 · 2015
Earlier work this paper cites.
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Cerebella: Automatic Generation of Nonverbal Behavior for Virtual Humans. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI ’15, 1)
Margot Lhommet, Yuyu Xu, and Stacy Marsella. 2015 · 2015
Earlier work this paper cites.
Learning from Gesture: How Our Hands Change Our Minds
Miriam Novack and Susan Goldin-Meadow. 2015 · 2015
Earlier work this paper cites.
Deep Unsupervised Learning Using Nonequilibrium Thermodynamics. In Proceedings of the International Conference on Machine Learning (ICML ’15) . 2256–2265
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015 · 2015
Earlier work this paper cites.
Joint Beat and Downbeat Tracking with Recurrent Neural Networks. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR ’16) . 255–261
Sebastian Böck, Florian Krebs, and Gerhard Widmer. 2016 · 2016
Earlier work this paper cites.
Long Short-Term Memory-Networks for Machine Reading. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP ’16) . ACL, 551–561
Jianpeng Cheng, Li Dong, and Mirella Lapata. 2016 · 2016
Earlier work this paper cites.
Motion Matching and the Road to Next-Gen Animation. In Proceedings of the Game Developers Conference (GDC ’16)
Simon Clavet. 2016 · 2016
Earlier work this paper cites.
Generative Choreography Using Deep Learning. In International Conference on Computational Creativity (ICCC ’16) . ACC, 272–277
Luka Crnkovic-Friis and Louise Crnkovic-Friis. 2016 · 2016
Earlier work this paper cites.
PERFORM: Perceptual Approach for Adding OCEAN Personality to Human Motion Using Laban Movement Analysis
Funda Durupinar, Mubbasir Kapadia, Susan Deutsch, Michael Neff, and Norman I. Badler. 2016 · 2016
Earlier work this paper cites.
Perception of Emotions and Body Movement in the Emilya Database
Nesrine Fourati and Catherine Pelachaud. 2016 · 2016
Earlier work this paper cites.
Minimum Entropy Rate Simplification of Stochastic Processes
Gustav Eje Henter and W. Bastiaan Kleijn. 2016 · 2016
Earlier work this paper cites.
A Deep Learning Framework for Character Motion Synthesis and Editing
Daniel Holden, Jun Saito, and Taku Komura. 2016 · 2016
Earlier work this paper cites.
A Note on the Evaluation of Generative Models
Lucas Theis, Aäron van den Oord, and Matthias Bethge. 2016 · 2016
Earlier work this paper cites.
WaveNet: A Generative Model for Raw Audio
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. 2016 · 2016
Earlier work this paper cites.
A Recurrent Variational Autoencoder for Human Motion Synthesis. In Proceedings of the British Machine Vision Conference (BMVC ’17) . BMVA Press, Article 119, 12 pages
Ikhansul Habibie, Daniel Holden, Jonathan Schwarz, Joe Yearsley, and Taku Komura. 2017 · 2017
Earlier work this paper cites.
GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium. In Advances in Neural Information Processing Systems (NIPS ’17)
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017 · 2017
Earlier work this paper cites.
Understanding the Impact of Animated Gesture Performance on Personality Perceptions
Harrison Jesse Smith and Michael Neff. 2017 · 2017
Earlier work this paper cites.
Speech-to-Gesture Generation: A Challenge in Deep Learning Approach with Bi-Directional LSTM. In Proceedings of the International Conference on Human Agent Interaction (HAI ’17) . ACM
Kenta Takeuchi, Dai Hasegawa, Shinichi Shirakawa, Naoshi Kaneko, Hiroshi Sakuta, and Kazuhiko Sumi. 2017 · 2017
Earlier work this paper cites.
Attention Is All You Need. In Advances in Neural Information Processing Systems (NIPS ’17) . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Investigating the Use of Recurrent Motion Modelling for Speech Gesture Generation. In Proceedings of the ACM International Conference on Intelligent Virtual Agents (IVA ’18) . ACM, 93–98
Ylva Ferstl and Rachel McDonnell. 2018 · 2018
Earlier work this paper cites.
Evaluation of Speech-to-Gesture Generation Using Bi-Directional LSTM Network. In Proceedings of the ACM International Conference on Intelligent Virtual Agents (IVA ’18) . ACM, 79–86
Dai Hasegawa, Naoshi Kaneko, Shinichi Shirakawa, Hiroshi Sakuta, and Kazuhiko Sumi. 2018 · 2018
Earlier work this paper cites.
Do WaveNets Dream of Acoustic Waves?
Kanru Hua. 2018 · 2018
Earlier work this paper cites.
FiLM: Visual Reasoning with a General Conditioning Layer. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI ’18, 1)
Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville. 2018 · 2018
Earlier work this paper cites.
Dance with Melody: An LSTM-Autoencoder Approach to Music-Oriented Dance Synthesis. In Proceedings of the ACM International Conference on Multimedia (MM ’18) . ACM, 1598–1606
Taoran Tang, Jia Jia, and Hanyang Mao. 2018 · 2018
Cited alongside, same era.
VAE with a VampPrior. In Proceedings of the International Conference on Artificial Intelligence and Statistics (AISTATS ’18) . 1214–1223
Jakub Tomczak and Max Welling. 2018 · 2018
Cited alongside, same era.
Mode-Adaptive Neural Networks for Quadruped Motion Control
He Zhang, Sebastian Starke, Taku Komura, and Jun Saito. 2018 · 2018
Cited alongside, same era.
What Do We Express Without Knowing? Emotion in Gesture. In Proceedings of the International Conference on Autonomous Agents and Multiagent Systems (AAMAS ’19) . IFAAMAS, 702–710
Gabriel Castillo and Michael Neff. 2019 · 2019
Cited alongside, same era.
BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics (NAACL ’19) . ACL, 4171–4186
Rhythm Is a Dancer: Music-Driven Motion Synthesis with Global Structure
Andreas Aristidou, Anastasios Yiannakidis, Kfir Aberman, Daniel Cohen-Or, Ariel Shamir, and Yiorgos Chrysanthou. 2022 · 2022
Closest in time.
eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers
Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, Bryan Catanzaro, et al · 2022
Closest in time.
Rishabh Dabral, Muhammad Hamza Mughal, Vladislav Golyanik, and Christian Theobalt. 2022 · 2022
Closest in time.
Guidance: A Cheat Code for Diffusion Models
Sander Dieleman. 2022 · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Multi-Objective Adversarial Gesture Generation. In Proceedings of the ACM SIGGRAPH Conference on Motion, Interaction and Games (MIG’19) . ACM, Article 3, 10 pages
Ylva Ferstl, Michael Neff, and Rachel McDonnell. 2019 · 2019
Cited alongside, same era.
Analyzing Input and Output Representations for Speech-Driven Gesture Generation. In Proceedings of the ACM International Conference on Intelligent Virtual Agents (IVA ’19) . ACM, 97–104
Taras Kucherenko, Dai Hasegawa, Gustav Eje Henter, Naoshi Kaneko, and Hedvig Kjellström. 2019 · 2019
Cited alongside, same era.
Talking With Hands 16.2M: A Large-Scale Dataset of Synchronized Body-Finger Motion and Audio for Conversational Motion Analysis and Synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV ’19) . 763–772
Gilwoo Lee, Zhiwei Deng, Shugao Ma, Takaaki Shiratori, Siddhartha Srinivasa, and Yaser Sheikh. 2019a · 2019
Cited alongside, same era.
Dancing to Music
Hsin-Ying Lee, Xiaodong Yang, Ming-Yu Liu, Ting-Chun Wang, Yu-Ding Lu, Ming-Hsuan Yang, and Jan Kautz. 2019b · 2019
Cited alongside, same era.
Speech-Driven Animation with Meaningful Behaviors
Najmeh Sadoughi and Carlos Busso. 2019 · 2019
Cited alongside, same era.
Generative Modeling by Estimating Gradients of the Data Distribution. In Advances in Neural Information Processing Systems (NeurIPS ’19)
Yang Song and Stefano Ermon. 2019 · 2019
Cited alongside, same era.
No Gestures Left Behind: Learning Relationships Between Spoken Language and Freeform Gestures. In Findings of the Association for Computational Linguistics (EMNLP ’20) . ACL, 1884–1895
Chaitanya Ahuja, Dong Won Lee, Ryo Ishii, and Louis-Philippe Morency. 2020a · 2020
Cited alongside, same era.
Mireille Fares, Michele Grimaldi, Catherine Pelachaud, and Nicolas Obin. 2022 · 2022
Closest in time.
Edmund J. C. Findlay, Haozheng Zhang, Ziyi Chang, and Hubert P. H. Shum. 2022 · 2022
Closest in time.
Exemplar-Based Stylized Gesture Generation from Speech: An Entry to the GENEA Challenge 2022. In Proceedings of the ACM International Conference on Multimodal Interaction (ICMI ’22) . ACM, 778–783
Saeed Ghorbani, Ylva Ferstl, and Marc-André Carbonneau. 2022 · 2022
Closest in time.
TM2T: Stochastic and Tokenized Modeling for the Reciprocal Generation of 3D Human Motions and Texts. In Proceedings of the European Conference on Computer Vision (ECCV ’22) . 580–597
Chuan Guo, Xinxin Zuo, Sen Wang, and Li Cheng. 2022 · 2022
Closest in time.
A Motion Matching-Based Framework for Controllable Gesture Synthesis from Speech. In ACM Special Interest Group on Computer Graphics and Interactive Techniques Conference Proceedings (SIGGRAPH ’22) . ACM, Article 46, 9 pages
Ikhsanul Habibie, Mohamed Elgharib, Kripasindhu Sarkar, Ahsan Abdullah, Simbarashe Nyatsanga, Michael Neff, and Christian Theobalt. 2022 · 2022
Closest in time.
Imagen Video: High Definition Video Generation with Diffusion Models
Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P. Kingma, Ben Poole, Mohammad Norouzi, David J. Fleet, et al · 2022
Closest in time.
Video Diffusion Models. In Advances in Neural Information Processing Systems (NeurIPS ’22) . 8633–8646
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J. Fleet. 2022b · 2022
Closest in time.
DiffPose: Multi-Hypothesis Human Pose Estimation Using Diffusion Models
Karl Holmquist and Bastian Wandt. 2022 · 2022
Closest in time.
Diffusion Models for Video Prediction and Infilling
Tobias Höppe, Arash Mehrjou, Stefan Bauer, Didrik Nielsen, and Andrea Dittadi. 2022 · 2022
Closest in time.
Multi-Scale Cascaded Generator for Music-Driven Dance Synthesis. In Proceedings of the International Joint Conference on Neural Networks (IJCNN ’22) . IEEE, 1–7
Hao Hu, Changhong Liu, Yong Chen, Aiwen Jiang, Zhenchun Lei, and Mingwen Wang. 2022 · 2022
Closest in time.
FLAME: Free-Form Language-Based Motion Synthesis & Editing
Jihoon Kim, Jiseob Kim, and Sungjoon Choi. 2022 · 2022
Closest in time.
Multimodal Analysis of the Predictability of Hand-Gesture Properties. In Proceedings of the International Conference on Autonomous Agents and Multiagent Systems (AAMAS ’22) . IFAAMAS, 770–779
Taras Kucherenko, Rajmund Nagy, Michael Neff, Hedvig Kjellström, and Gustav Eje Henter. 2022 · 2022
Closest in time.
BDDM: Bilateral Denoising Diffusion Models for Fast and High-Quality Speech Synthesis. In Proceedings of the International Conference on Learning Representations (ICLR ’22)
Max W. Y. Lam, Jun Wang, Dan Su, and Dong Yu. 2022 · 2022
Closest in time.
DanceFormer: Music Conditioned 3D Dance Generation with Parametric Motion Transformer. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI ’22, Vol. 36) . 1272–1279
Buyu Li, Yongchi Zhao, Shi Zhelun, and Lu Sheng. 2022 · 2022
Closest in time.
Pretrained Diffusion Models for Unified Human Motion Synthesis
Jianxin Ma, Shuai Bai, and Chang Zhou. 2022 · 2022
Closest in time.
Real-Time Style Modelling of Human Locomotion via Feature-Wise Transformations and Local Motion Phases
Ian Mason, Sebastian Starke, and Taku Komura. 2022 · 2022
Closest in time.
On Distillation of Guided Diffusion Models. In Proceedings of the NeurIPS Workshop on Score-Based Methods (NeurIPS ’22 Workshop)
Chenlin Meng, Ruiqi Gao, Diederik P. Kingma, Stefano Ermon, Jonathan Ho, and Tim Salimans. 2022 · 2022
Closest in time.
GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models. In Proceedings of the International Conference on Machine Learning (ICML -22) . 16784–16804
Alexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob Mcgrew, Ilya Sutskever, and Mark Chen. 2022 · 2022
Closest in time.
TEMOS: Generating Diverse Human Motions from Textual Descriptions. In Proceedings of the European Conference on Computer Vision (ECCV ’22) . Springer, 480–497
Mathis Petrovich, Michael J. Black, and Gül Varol. 2022 · 2022
Closest in time.
Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation. In Proceedings of the International Conference on Learning Representations (ICLR ’22)
Ofir Press, Noah Smith, and Mike Lewis. 2022 · 2022
Closest in time.
Hierarchical Text-Conditional Image Generation with CLIP Latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022 · 2022
Closest in time.
High-Resolution Image Synthesis with Latent Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR ’22) . 10684–10695
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Closest in time.
Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding. In Advances in Neural Information Processing Systems (NeurIPS ’22) . 36479–36494
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, et al · 2022
Closest in time.
Progressive Distillation for Fast Sampling of Diffusion Models. In Proceedings of the International Conference on Learning Representations (ICLR ’22)
Tim Salimans and Jonathan Ho. 2022 · 2022
Closest in time.
Bailando: 3D Dance Generation by Actor-Critic GPT with Choreographic Memory. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR ’22) . 11050–11059
Li Siyao, Weijiang Yu, Tianpei Gu, Chunze Lin, Quan Wang, Chen Qian, Chen Change Loy, and Ziwei Liu. 2022 · 2022
Closest in time.
DeepPhase: Periodic Autoencoders for Learning Motion Phase Manifolds
Sebastian Starke, Ian Mason, and Taku Komura. 2022 · 2022
Closest in time.
MCVD – Masked Conditional Video Diffusion for Prediction, Generation, and Interpolation. In Advances in Neural Information Processing Systems (NeurIPS ’22)
Vikram Voleti, Alexia Jolicoeur-Martineau, and Christopher Pal. 2022 · 2022
Closest in time.
Learning Soccer Juggling Skills with Layer-Wise Mixture-of-Experts. In ACM Special Interest Group on Computer Graphics and Interactive Techniques Conference Proceedings (SIGGRAPH ’22) . ACM, Article 25, 9 pages
Zhaoming Xie, Sebastian Starke, Hung Yu Ling, and Michiel van de Panne. 2022 · 2022
Closest in time.
Gesture2Vec: Clustering Gestures Using Representation Learning Methods for Co-Speech Gesture Generation. In Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS ’22) . IEEE
Payam Jome Yazdian, Mo Chen, and Angelica Lim. 2022 · 2022
Closest in time.
The GENEA Challenge 2022: A Large Evaluation of Data-Driven Co-Speech Gesture Generation. In Proceedings of the ACM International Conference on Multimodal Interaction (ICMI ’22) . ACM, 736–747
Youngwoo Yoon, Pieter Wolfert, Taras Kucherenko, Carla Viegas, Teodor Nikolov, Mihail Tsakov, and Gustav Eje Henter. 2022 · 2022
Closest in time.
MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model
Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, and Ziwei Liu. 2022a · 2022
Closest in time.
EGSDE: Unpaired Image-to-Image Translation via Energy-Guided Stochastic Differential Equations. In Advances in Neural Information Processing Systems (NeurIPS ’22) . 3609–3623
Min Zhao, Fan Bao, Chongxuan Li, and Jun Zhu. 2022 · 2022
Closest in time.
GestureMaster: Graph-Based Speech-Driven Gesture Generation. In Proceedings of the ACM International Conference on Multimodal Interaction (ICMI ’22) . ACM, 764–770
Chi Zhou, Tengyue Bian, and Kang Chen. 2022 · 2022
Closest in time.
MotionBERT: Unified Pretraining for Human Motion Analysis
Wentao Zhu, Xiaoxuan Ma, Zhaoyang Liu, Libin Liu, Wayne Wu, and Yizhou Wang. 2022 · 2022
Closest in time.
Music2Dance: DanceNet for Music-Driven Dance Generation
Wenlin Zhuang, Congyi Wang, Jinxiang Chai, Yangang Wang, Ming Shao, and Siyu Xia. 2022 · 2022
Closest in time.
GestureDiffuCLIP: Gesture Diffusion Model with CLIP Latents
Tenglong Ao, Zeyi Zhang, and Libin Liu. 2023 · 2023
Closest in time.
ZeroEGGS: Zero-Shot Example-Based Gesture Generation from Speech
Saeed Ghorbani, Ylva Ferstl, Daniel Holden, Nikolaus F. Troje, and Marc-André Carbonneau. 2023 · 2023
Closest in time.
Evaluating Gesture-Generation in a Large-Scale Open Challenge: The GENEA Challenge 2022
Taras Kucherenko, Pieter Wolfert, Youngwoo Yoon, Carla Viegas, Teodor Nikolov, Mihail Tsakov, and Gustav Eje Henter. 2023 · 2023
Closest in time.
A Comprehensive Review of Data-Driven Co-Speech Gesture Generation
Simbarashe Nyatsanga, Taras Kucherenko, Chaitanya Ahuja, Gustav Eje Henter, and Michael Neff. 2023 · 2023
Closest in time.
PirouNet: Creating Dance Through Artist-Centric Deep Learning. In Proceedings of the EAI International Conference ArtsIT, Interactivity and Game Creation (ArtsIT ’23) . Springer Nature, 447–465
Mathilde Papillon, Mariel Pettee, and Nina Miolane. 2023 · 2023
Closest in time.
Human Motion Diffusion Model. In Proceedings of the International Conference on Learning Representations (ICLR ’23)
Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Daniel Cohen-Or, and Amit H. Bermano. 2023 · 2023
Closest in time.
Jonathan Tseng, Rodrigo Castellon, and C. Karen Liu. 2023 · 2023
Closest in time.
Dance Style Transfer with Cross-Modal Transformer. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV ’23) . 5047–5056
Wenjie Yin, Hang Yin, Kim Baraka, Danica Kragic, and Mårten Björkman. 2023 · 2023
Closest in time.
DiffMotion: Speech-Driven Gesture Synthesis Using Denoising Diffusion Model. In Proceedings of the International Conference on Multimedia Modeling (MMM ’23) . 231–242
Fan Zhang, Naye Ji, Fuxing Gao, and Yongping Li. 2023 · 2023
Closest in time.