Fetching the paper…
Reading the bibliography…
Generating the motion of orchestral conductors from a given piece of symphony music is a challenging task since it requires a model to learn semantic music features and capture the underlying distribution of real conducting motion.
Music-oriented dance video synthesis with pose perceptual loss
Ren, X.; Li, H.; Huang, Z.; and Chen, Q. 2019 · 1912
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J.; Meng, C.; and Ermon, S. 2020 · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Earlier work this paper cites.
librosa: Audio and music signal analysis in python
McFee, B.; Raffel, C.; Liang, D.; Ellis, D. P.; McVicar, M.; Battenberg, E.; and Nieto, O. 2015 · 2015
Earlier work this paper cites.
Sequential deep learning for dancing motion generation
Yalta, N.; Ogata, T.; and Nakadai, K. 2016 · 2016
Earlier work this paper cites.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L. u.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Listen to dance: Music-driven choreography generation using autoregressive encoder-decoder network
Lee, J.; Kim, S.; and Lee, K. 2018 · 2018
Earlier work this paper cites.
Skeleton Plays Piano: Online Generation of Pianist Body Movements from MIDI Performance
Li, B.; Maezawa, A.; and Duan, Z. 2018 · 2018
Earlier work this paper cites.
Dance with melody: An lstm-autoencoder approach to music-oriented dance synthesis
Tang, T.; Jia, J.; and Mao, H. 2018 · 2018
Earlier work this paper cites.
Spatial temporal graph convolutional networks for skeleton-based action recognition
Yan, S.; Xiong, Y.; and Lin, D. 2018 · 2018
Earlier work this paper cites.
Robots learn social skills: End-to-end learning of co-speech gesture generation for humanoid robots
Yoon, Y.; Ko, W.-R.; Jang, M.; Lee, J.; Kim, J.; and Lee, G. 2019 · 2019
Earlier work this paper cites.
No gestures left behind: Learning relationships between spoken language and freeform gestures
Ahuja, C.; Lee, D. W.; Ishii, R.; and Morency, L.-P. 2020 · 2020
Cited alongside, same era.
Generative adversarial networks
Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; and Bengio, Y. 2020 · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Cited alongside, same era.
Temporally guided music-to-body-movement generation
Kao, H.-K.; and Su, L. 2020 · 2020
Cited alongside, same era.
DeepDance: music-to-dance motion choreography with adversarial learning
Sun, G.; Wong, Y.; Cheng, Z.; Kankanhalli, M. S.; Geng, W.; and Li, X. 2020 · 2020
Cited alongside, same era.
Speech gesture generation from the trimodal context of text, audio, and speaker identity
Yoon, Y.; Cha, B.; Lee, J.-H.; Jang, M.; Lee, J.; Kim, J.; and Lee, G. 2020 · 2020
Speech drives templates: Co-speech gesture synthesis with learned templates
Qian, S.; Tu, Z.; Zhi, Y.; Liu, W.; and Gao, S. 2021 · 2021
Later among the works it cites.
High-Resolution Image Synthesis with Latent Diffusion Models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2021 · 2021
Later among the works it cites.
Classifier-free diffusion guidance
Ho, J.; and Salimans, T. 2022 · 2022
Later among the works it cites.
A Survey on Deep Learning for Skeleton-Based Human Animation
Mourot, L.; Hoyet, L.; Le Clerc, F.; Schnitzler, F.; and Hellier, P. 2022 · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Random erasing data augmentation
Zhong, Z.; Zheng, L.; Kang, G.; Li, S.; and Yang, Y. 2020 · 2020
Cited alongside, same era.
VirtualConductor: Music-driven Conducting Video Generation System
Chen, D.; Liu, F.; Li, Z.; and Xu, F. 2021 · 2021
Cited alongside, same era.
Diffusion models beat gans on image synthesis
Dhariwal, P.; and Nichol, A. 2021 · 2021
Cited alongside, same era.
Ai choreographer: Music conditioned 3d dance generation with aist++
Li, R.; Yang, S.; Ross, D. A.; and Kanazawa, A. 2021 · 2021
Cited alongside, same era.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Nichol, A.; Dhariwal, P.; Ramesh, A.; Shyam, P.; Mishkin, P.; McGrew, B.; Sutskever, I.; and Chen, M. 2021 · 2021
Cited alongside, same era.
Self-supervised music motion synchronization learning for music-driven conducting motion generation
Liu, F.; Chen, D.-L.; Zhou, R.-Z.; Yang, S.; and Xu, F. 2022a
Cited in the paper.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E.; Ghasemipour, S. K. S.; Ayan, B. K.; Mahdavi, S. S.; Lopes, R. G.; et al. 2022 · 2022
Later among the works it cites.
Tevet, G.; Raab, S.; Gordon, B.; Shafir, Y.; Cohen-Or, D.; and Bermano, A. H. 2022 · 2022
Later among the works it cites.
MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model
Zhang, M.; Cai, Z.; Pan, L.; Hong, F.; Guo, X.; Yang, L.; and Liu, Z. 2022 · 2022
Later among the works it cites.
Taming Diffusion Models for Audio-Driven Co-Speech Gesture Generation
Zhu, L.; Liu, X.; Liu, X.; Qian, R.; Liu, Z.; and Yu, L. 2023 · 2023
Closest in time.