Fetching the paper…
Reading the bibliography…
The automatic generation of stylized co-speech gestures has recently received increasing attention.
Hand and Mind: What Gestures Reveal about Thought
David McNeill. 1992 · 1992
Earlier work this paper cites.
Animated Conversation: Rule-Based Generation of Facial Expression, Gesture & Spoken Intonation for Multiple Conversational Agents. In Proceedings of the 21st Annual Conference on Computer Graphics and Interactive Techniques (SIGGRAPH ’94) . Association for Computing Machinery, New York, NY, USA, 413–420
Justine Cassell, Catherine Pelachaud, Norman Badler, Mark Steedman, Brett Achorn, Tripp Becket, Brett Douville, Scott Prevost, and Matthew Stone. 1994 · 1994
Earlier work this paper cites.
BEAT: The Behavior Expression Animation Toolkit. In Proceedings of the 28th Annual Conference on Computer Graphics and Interactive Techniques (SIGGRAPH ’01) . Association for Computing Machinery, New York, NY, USA, 477–486
Justine Cassell, Hannes Högni Vilhjálmsson, and Timothy Bickmore. 2001 · 2001
Earlier work this paper cites.
Gesture Generation by Imitation: From Human Behavior to Computer Character Animation
Michael Kipp. 2004 · 2004
Earlier work this paper cites.
Style Translation for Human Motion
Eugene Hsu, Kari Pulli, and Jovan Popović. 2005 · 2005
Earlier work this paper cites.
Towards a Common Framework for Multimodal Generation: The Behavior Markup Language. In Proceedings of the 6th International Conference on Intelligent Virtual Agents (Marina Del Rey, CA) (IVA’06) . Springer-Verlag, Berlin, Heidelberg, 205–217
Stefan Kopp, Brigitte Krenn, Stacy Marsella, Andrew N. Marshall, Catherine Pelachaud, Hannes Pirker, Kristinn R. Thórisson, and Hannes Vilhjálmsson. 2006 · 2006
Earlier work this paper cites.
Gesture Modeling and Animation Based on a Probabilistic Re-Creation of Speaker Style
Michael Neff, Michael Kipp, Irene Albrecht, and Hans-Peter Seidel. 2008 · 2008
Earlier work this paper cites.
Real-Time Prosody-Driven Synthesis of Body Language
Sergey Levine, Christian Theobalt, and Vladlen Koltun. 2009 · 2009
Earlier work this paper cites.
Gesture Controllers
Sergey Levine, Philipp Krähenbühl, Sebastian Thrun, and Vladlen Koltun. 2010 · 2010
Earlier work this paper cites.
Modeling Style and Variation in Human Motion. In Proceedings of the 2010 ACM SIGGRAPH/Eurographics Symposium on Computer Animation (Madrid, Spain) (SCA ’10) . Eurographics Association, Goslar, DEU, 21–30
Wanli Ma, Shihong Xia, Jessica K. Hodgins, Xiao Yang, Chunpeng Li, and Zhaoqi Wang. 2010 · 2010
Earlier work this paper cites.
Gesture and speech in interaction: An overview
Petra Wagner, Zofia Malisz, and Stefan Kopp. 2014 · 2013
Earlier work this paper cites.
SMPL: A Skinned Multi-Person Linear Model
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black. 2015 · 2015
Earlier work this paper cites.
Deep Unsupervised Learning using Nonequilibrium Thermodynamics. In Proceedings of the 32nd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 37) , Francis Bach and David Blei (Eds.). PMLR, Lille, France, 2256–2265
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015 · 2015
Earlier work this paper cites.
Realtime Style Transfer for Unlabeled Heterogeneous Human Motion
Shihong Xia, Congyi Wang, Jinxiang Chai, and Jessica Hodgins. 2015 · 2015
Earlier work this paper cites.
A Deep Learning Framework for Character Motion Synthesis and Editing
Daniel Holden, Jun Saito, and Taku Komura. 2016 · 2016
Earlier work this paper cites.
Spectral Style Transfer for Human Motion between Independent Actions
M. Ersin Yumer and Niloy J. Mitra. 2016 · 2016
Earlier work this paper cites.
Phase-Functioned Neural Networks for Character Control
Daniel Holden, Taku Komura, and Jun Saito. 2017 · 2017
Earlier work this paper cites.
Arbitrary Style Transfer in Real-Time With Adaptive Instance Normalization. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)
Xun Huang and Serge Belongie. 2017 · 2017
Earlier work this paper cites.
Montreal Forced Aligner: Trainable Text-Speech Alignment Using Kaldi.. In Interspeech , Vol. 2017. 498–502
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger. 2017 · 2017
Earlier work this paper cites.
Neural Discrete Representation Learning. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, USA) (NIPS’17) . Curran Associates Inc., Red Hook, NY, USA, 6309–6318
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need. In Advances in Neural Information Processing Systems , I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30. Curran Associates, Inc
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Few-shot Learning of Homogeneous Human Locomotion Styles
I. Mason, S. Starke, H. Zhang, H. Bilen, and T. Komura. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . Association for Computational Linguistics, Minneapolis, Minnesota, 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Stylistic Locomotion Modeling and Synthesis Using Variational Generative Models. In Proceedings of the 12th ACM SIGGRAPH Conference on Motion, Interaction and Games (Newcastle upon Tyne, United Kingdom) (MIG ’19) . Association for Computing Machinery, New York, NY, USA, Article 32, 10 pages
Han Du, Erik Herrmann, Janis Sprenger, Klaus Fischer, and Philipp Slusallek. 2019 · 2019
Earlier work this paper cites.
Learning Individual Styles of Conversational Gesture. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
Shiry Ginosar, Amir Bar, Gefen Kohavi, Caroline Chan, Andrew Owens, and Jitendra Malik. 2019 · 2019
Earlier work this paper cites.
Robots Learn Social Skills: End-to-End Learning of Co-Speech Gesture Generation for Humanoid Robots. In 2019 International Conference on Robotics and Automation (ICRA) . 4303–4309
Youngwoo Yoon, Woo-Ri Ko, Minsu Jang, Jaeyeon Lee, Jaehong Kim, and Geehyuk Lee. 2019 · 2019
Earlier work this paper cites.
Yi Zhou, Connelly Barnes, Jingwan Lu, Jimei Yang, and Hao Li. 2019 · 2019
Earlier work this paper cites.
Unpaired Motion Style Transfer from Video to Animation
Kfir Aberman, Yijia Weng, Dani Lischinski, Daniel Cohen-Or, and Baoquan Chen. 2020 · 2020
Cited alongside, same era.
Style Transfer for Co-speech Gesture Animation: A Multi-speaker Conditional-Mixture Approach. In Computer Vision – ECCV 2020 , Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm (Eds.). Springer International Publishing, Cham, 248–265
Chaitanya Ahuja, Dong Won Lee, Yukiko I. Nakano, and Louis-Philippe Morency. 2020 · 2020
Cited alongside, same era.
Style-Controllable Speech-Driven Gesture Synthesis Using Normalising Flows
Simon Alexanderson, Gustav Eje Henter, Taras Kucherenko, and Jonas Beskow. 2020 · 2020
Cited alongside, same era.
Jukebox: A generative model for music
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever. 2020 · 2020
Cited alongside, same era.
GANimator: Neural Motion Synthesis from a Single Sequence
Peizhuo Li, Kfir Aberman, Zihan Zhang, Rana Hanocka, and Olga Sorkine-Hornung. 2022 · 2022
Later among the works it cites.
SEEG: Semantic Energized Co-Speech Gesture Generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 10473–10482
Yuanzhi Liang, Qianyu Feng, Linchao Zhu, Li Hu, Pan Pan, and Yi Yang. 2022 · 2022
Later among the works it cites.
Murf.AI: An Online Text-to-Speech Tool
Murf.AI. 2022 · 2022
Later among the works it cites.
ChatGPT: Optimizing Language Models for Dialogue
OpenAI. 2022 · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Denoising Diffusion Probabilistic Models. In Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 6840–6851
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Cited alongside, same era.
Gesticulator: A Framework for Semantically-Aware Speech-Driven Gesture Generation. In Proceedings of the 2020 International Conference on Multimodal Interaction (Virtual Event, Netherlands) (ICMI ’20) . Association for Computing Machinery, New York, NY, USA, 242–250
Taras Kucherenko, Patrik Jonell, Sanne van Waveren, Gustav Eje Henter, Simon Alexandersson, Iolanda Leite, and Hedvig Kjellström. 2020 · 2020
Cited alongside, same era.
Improved Techniques for Training Score-Based Generative Models. In Proceedings of the 34th International Conference on Neural Information Processing Systems (Vancouver, BC, Canada) (NIPS’20) . Curran Associates Inc., Red Hook, NY, USA, Article 1043, 11 pages
Yang Song and Stefano Ermon. 2020 · 2020
Cited alongside, same era.
Speech Gesture Generation from the Trimodal Context of Text, Audio, and Speaker Identity
Youngwoo Yoon, Bok Cha, Joo-Haeng Lee, Minsu Jang, Jaeyeon Lee, Jaehong Kim, and Geehyuk Lee. 2020 · 2020
Cited alongside, same era.
Text2Gestures: A Transformer-Based Network for Generating Emotive Body Gestures for Virtual Agents. In 2021 IEEE Conference on Virtual Reality and 3D User Interfaces (IEEE VR) . IEEE
Uttaran Bhattacharya, Nicholas Rewkowski, Abhishek Banerjee, Pooja Guhan, Aniket Bera, and Dinesh Manocha. 2021b · 2021
Cited alongside, same era.
Learning Speech-Driven 3D Conversational Gestures from Video. In Proceedings of the 21st ACM International Conference on Intelligent Virtual Agents (Virtual Event, Japan) (IVA ’21) . Association for Computing Machinery, New York, NY, USA, 101–108
Ikhsanul Habibie, Weipeng Xu, Dushyant Mehta, Lingjie Liu, Hans-Peter Seidel, Gerard Pons-Moll, Mohamed Elgharib, and Christian Theobalt. 2021 · 2021
Cited alongside, same era.
Classifier-Free Diffusion Guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications
Jonathan Ho and Tim Salimans. 2021 · 2021
Cited alongside, same era.
Perceiver: General Perception with Iterative Attention. In Proceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 139) , Marina Meila and Tong Zhang (Eds.). PMLR, 4651–4664
Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Joao Carreira. 2021 · 2021
Cited alongside, same era.
High-Resolution Image Synthesis With Latent Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 10684–10695
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Later among the works it cites.
Motionclip: Exposing human motion generation to clip space. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXII . Springer, 358–374
Guy Tevet, Brian Gordon, Amir Hertz, Amit H Bermano, and Daniel Cohen-Or. 2022 · 2022
Later among the works it cites.
Gesture2Vec: Clustering Gestures using Representation Learning Methods for Co-speech Gesture Generation. In 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . 3100–3107
Payam Jome Yazdian, Mo Chen, and Angelica Lim. 2022 · 2022
Later among the works it cites.
Audio-Driven Stylized Gesture Generation with Flow-Based Model. In Computer Vision – ECCV 2022 , Shai Avidan, Gabriel Brostow, Moustapha Cissé, Giovanni Maria Farinella, and Tal Hassner (Eds.). Springer Nature Switzerland, Cham, 712–728
Sheng Ye, Yu-Hui Wen, Yanan Sun, Ying He, Ziyang Zhang, Yaoyuan Wang, Weihua He, and Yong-Jin Liu. 2022 · 2022
Later among the works it cites.
Generating Holistic 3D Human Motion from Speech
Hongwei Yi, Hualin Liang, Yifei Liu, Qiong Cao, Yandong Wen, Timo Bolkart, Dacheng Tao, and Michael J. Black. 2022 · 2022
Later among the works it cites.
The GENEA Challenge 2022: A Large Evaluation of Data-Driven Co-Speech Gesture Generation. In Proceedings of the 2022 International Conference on Multimodal Interaction (Bengaluru, India) (ICMI ’22) . Association for Computing Machinery, New York, NY, USA, 736–747
Youngwoo Yoon, Pieter Wolfert, Taras Kucherenko, Carla Viegas, Teodor Nikolov, Mihail Tsakov, and Gustav Eje Henter. 2022 · 2022
Later among the works it cites.
MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model
Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, and Ziwei Liu. 2022 · 2022
Later among the works it cites.
GestureMaster: Graph-Based Speech-Driven Gesture Generation. In Proceedings of the 2022 International Conference on Multimodal Interaction (Bengaluru, India) (ICMI ’22) . Association for Computing Machinery, New York, NY, USA, 764–770
Chi Zhou, Tengyue Bian, and Kang Chen. 2022 · 2022
Later among the works it cites.
Fireplace 10 hours full HD
Fireplace10hours. 2016 · 2023
Closest in time.
ZeroEGGS: Zero-shot Example-based Gesture Generation from Speech
Saeed Ghorbani, Ylva Ferstl, Daniel Holden, Nikolaus F. Troje, and Marc-André Carbonneau. 2023 · 2023
Closest in time.
Wiz Khalifa - Roll Up
Wiz Khalifa. 2011 · 2023
Closest in time.
Lightning before the thunder
manabouttown1. 2021 · 2023
Closest in time.
A Comprehensive Review of Data-Driven Co-Speech Gesture Generation
Simbarashe Nyatsanga, Taras Kucherenko, Chaitanya Ahuja, Gustav Eje Henter, and Michael Neff. 2023 · 2023
Closest in time.
Swaying Trees in The Wind, Rumbling Leaves, Relaxing Wind
RelaxingSoundsOfNature. 2018 · 2023
Closest in time.
Human Motion Diffusion Model. In The Eleventh International Conference on Learning Representations
Guy Tevet, Sigal Raab, Brian Gordon, Yoni Shafir, Daniel Cohen-or, and Amit Haim Bermano. 2023 · 2023
Closest in time.
BIRDS IN FLIGHT
Wildlife_World. 2019 · 2023
Closest in time.
Jurassic World Evolution - All 48 Dinosaurs (1080p 60FPS)
worldofdinosaurs1. 2018 · 2023
Closest in time.
Yoga Party | 30-Minute Home Yoga Practice
yogawithadriene. 2020 · 2023
Closest in time.
Center - Day 29 - Pleasure
yogawithadriene. 2023 · 2023
Closest in time.
DiffMotion: Speech-Driven Gesture Synthesis Using Denoising Diffusion Model. In MultiMedia Modeling: 29th International Conference, MMM 2023, Bergen, Norway, January 9–12, 2023, Proceedings, Part I . Springer, 231–242
Fan Zhang, Naye Ji, Fuxing Gao, and Yongping Li. 2023 · 2023
Closest in time.
Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression Learning. In Proceedings of the 29th ACM International Conference on Multimedia (Virtual Event, China) (MM ’21) . Association for Computing Machinery, New York, NY, USA, 2027–2036
Uttaran Bhattacharya, Elizabeth Childs, Nicholas Rewkowski, and Dinesh Manocha. 2021a · 2036
Closest in time.
StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . 2085–2094
Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, and Dani Lischinski. 2021 · 2094
Closest in time.