Fetching the paper…
Reading the bibliography…
Generating vivid and diverse 3D co-speech gestures is crucial for various applications in animating virtual avatars.
J. Cassell, C. Pelachaud, N. Badler, M. Steedman, B. Achorn, T. Becket, B. Douville, S. Prevost, and M. Stone, “Animated conversation: rule-based generation of facial expression, gesture & spoken intonation for multiple conversational agents,” in Proceedings of the 21st annual conference on Computer graphics and interactive techniques , 1994, pp. 413–420
1994
Earlier work this paper cites.
J. Cassell, D. McNeill, and K.-E. McCullough, “Speech-gesture mismatches: Evidence for one underlying representation of linguistic and nonlinguistic information,” Pragmatics & cognition , vol. 7, no. 1, pp. 1–34, 1999
1999
Earlier work this paper cites.
I. Poggi, C. Pelachaud, F. de Rosis, V. Carofiglio, and B. De Carolis, “Greta. a believable embodied conversational agent,” Multimodal intelligent information presentation , pp. 3–25, 2005
2005
Earlier work this paper cites.
J. P. Bello, L. Daudet, S. Abdallah, C. Duxbury, M. Davies, and M. B. Sandler, “A tutorial on onset detection in music signals,” IEEE Transactions on speech and audio processing , vol. 13, no. 5, pp. 1035–1047, 2005
2005
Earlier work this paper cites.
I. Akhter, Y. Sheikh, S. Khan, and T. Kanade, “Nonrigid structure from motion in trajectory space,” Advances in neural information processing systems , vol. 21, 2008
2008
Earlier work this paper cites.
J. P. De Ruiter, A. Bangerter, and P. Dings, “The interplay between gesture and speech in the production of referring expressions: Investigating the tradeoff hypothesis,” Topics in cognitive science , vol. 4, no. 2, pp. 232–248, 2012
2012
Earlier work this paper cites.
C.-M. Huang and B. Mutlu, “Robot behavior toolkit: generating effective social behaviors for robots,” in Proceedings of the seventh annual ACM/IEEE international conference on Human-Robot Interaction , 2012, pp. 25–32
2012
Earlier work this paper cites.
M. Salem, S. Kopp, I. Wachsmuth, K. Rohlfing, and F. Joublin, “Generation and evaluation of communicative robot gesture,” International Journal of Social Robotics , vol. 4, pp. 201–217, 2012
2012
Earlier work this paper cites.
S. Marsella, Y. Xu, M. Lhommet, A. Feng, S. Scherer, and A. Shapiro, “Virtual character performance from speech,” in Proceedings of the 12th ACM SIGGRAPH/Eurographics symposium on computer animation , 2013, pp. 25–35
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
P. Wagner, Z. Malisz, and S. Kopp, “Gesture and speech in interaction: An overview,” pp. 209–232, 2014
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
K. Sohn, H. Lee, and X. Yan, “Learning structured output representation using deep conditional generative models,” Advances in neural information processing systems , vol. 28, 2015
2015
Earlier work this paper cites.
S. Orts-Escolano, C. Rhemann, S. Fanello, W. Chang, A. Kowdle, Y. Degtyarev, D. Kim, P. L. Davidson, S. Khamis, M. Dou et al. , “Holoportation: Virtual 3d teleportation in real-time,” in Proceedings of the 29th annual symposium on user interface software and technology , 2016, pp. 741–754
2016
Earlier work this paper cites.
Y. Huang, F. Bogo, C. Lassner, A. Kanazawa, P. V. Gehler, J. Romero, I. Akhter, and M. J. Black, “Towards accurate marker-less human shape and pose estimation over time,” in 2017 international conference on 3D vision (3DV) . IEEE, 2017, pp. 421–430
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
C. T. Ishi, D. Machiyashiki, R. Mikata, and H. Ishiguro, “A speech-driven hand gesture generation method and evaluation in android robots,” IEEE Robotics and Automation Letters , vol. 3, no. 4, pp. 3757–3764, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Y. Yoon, W.-R. Ko, M. Jang, J. Lee, J. Kim, and G. Lee, “Robots learn social skills: End-to-end learning of co-speech gesture generation for humanoid robots,” in 2019 International Conference on Robotics and Automation (ICRA) . IEEE, 2019, pp. 4303–4309
2019
Earlier work this paper cites.
C. Ahuja and L.-P. Morency, “Language2pose: Natural language grounded pose forecasting,” in 2019 International Conference on 3D Vision (3DV) . IEEE, 2019, pp. 719–728
2019
Cited alongside, same era.
W. Mao, M. Liu, M. Salzmann, and H. Li, “Learning trajectory dependencies for human motion prediction,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 9489–9497
2019
Cited alongside, same era.
S. Ginosar, A. Bar, G. Kohavi, C. Chan, A. Owens, and J. Malik, “Learning individual styles of conversational gesture,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 3497–3506
2019
Cited alongside, same era.
Y. Zhou, C. Barnes, J. Lu, J. Yang, and H. Li, “On the continuity of rotation representations in neural networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 5745–5753
2019
Cited alongside, same era.
2022
Later among the works it cites.
T. Ao, Q. Gao, Y. Lou, B. Chen, and L. Liu, “Rhythmic gesticulator: Rhythm-aware co-speech gesture synthesis with hierarchical neural embeddings,” ACM Transactions on Graphics (TOG) , vol. 41, no. 6, pp. 1–19, 2022
2022
Later among the works it cites.
H. Liu, Z. Zhu, N. Iwamoto, Y. Peng, Z. Li, Y. Zhou, E. Bozkurt, and B. Zheng, “Beat: A large-scale semantic and emotional multi-modal dataset for conversational gestures synthesis,” in Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part VII . Springer, 2022, pp. 612–630
2022
Later among the works it cites.
A. Huang, P. Knierim, F. Chiossi, L. L. Chuang, and R. Welsch, “Proxemics for human-agent interaction in augmented reality,” in CHI Conference on Human Factors in Computing Systems , 2022, pp. 1–13
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H.-Y. Lee, X. Yang, M.-Y. Liu, T.-C. Wang, Y.-D. Lu, M.-H. Yang, and J. Kautz, “Dancing to music,” Advances in neural information processing systems , vol. 32, 2019
2019
Cited alongside, same era.
Y. Yoon, B. Cha, J.-H. Lee, M. Jang, J. Lee, J. Kim, and G. Lee, “Speech gesture generation from the trimodal context of text, audio, and speaker identity,” ACM Transactions on Graphics (TOG) , vol. 39, no. 6, pp. 1–16, 2020
2020
Cited alongside, same era.
V. Choutas, G. Pavlakos, T. Bolkart, D. Tzionas, and M. J. Black, “Monocular expressive body regression through body-driven attention,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part X 16 . Springer, 2020, pp. 20–40
2020
Cited alongside, same era.
2020
Cited alongside, same era.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Advances in Neural Information Processing Systems , vol. 33, pp. 6840–6851, 2020
2020
Cited alongside, same era.
J. Li, D. Kang, W. Pei, X. Zhe, Y. Zhang, Z. He, and L. Bao, “Audio2gestures: Generating diverse gestures from speech audio with conditional variational autoencoders,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 11 293–11 302
2021
Cited alongside, same era.
N. Battan, Y. Agrawal, S. S. Rao, A. Goel, and A. Sharma, “Glocalnet: Class-aware long-term human motion synthesis,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision , 2021, pp. 879–888
2021
Cited alongside, same era.
C. Zheng, S. Zhu, M. Mendieta, T. Yang, C. Chen, and Z. Ding, “3d human pose estimation with spatial and temporal transformers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 11 656–11 665
2021
Cited alongside, same era.
2022
Later among the works it cites.
Z. Wang, X. Qi, K. Yuan, and M. Sun, “Self-supervised correlation mining network for person image generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 7703–7712
2022
Later among the works it cites.
C. Zhong, L. Hu, Z. Zhang, Y. Ye, and S. Xia, “Spatio-temporal gating-adjacency gcn for human motion prediction,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 6447–6456
2022
Later among the works it cites.
W. Mao, R. I. Hartley, M. Salzmann et al. , “Contact-aware human motion forecasting,” Advances in Neural Information Processing Systems , vol. 35, pp. 7356–7367, 2022
2022
Later among the works it cites.
W. Mao, M. Liu, and M. Salzmann, “Weakly-supervised action transition learning for stochastic human motion prediction,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 8151–8160
2022
Later among the works it cites.
2022
Later among the works it cites.
R. Huang, W. Zhong, and G. Li, “Audio-driven talking head generation with transformer and 3d morphable model,” in Proceedings of the 30th ACM International Conference on Multimedia , 2022, pp. 7035–7039
2022
Later among the works it cites.
J. Jiang, P. Streli, H. Qiu, A. Fender, L. Laich, P. Snape, and C. Holz, “Avatarposer: Articulated full-body pose tracking from sparse motion sensing,” in Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part V . Springer, 2022, pp. 443–460
2022
Later among the works it cites.
L. Zhu, X. Liu, X. Liu, R. Qian, Z. Liu, and L. Yu, “Taming diffusion models for audio-driven co-speech gesture generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023
2023
Closest in time.
H. Yi, H. Liang, Y. Liu, Q. Cao, Y. Wen, T. Bolkart, D. Tao, and M. J. Black, “Generating holistic 3d human motion from speech,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023
2023
Closest in time.
S. Yang, Z. Wu, M. Li, Z. Zhang, L. Hao, W. Bao, M. Cheng, and L. Xiao, “Diffusestylegesture: Stylized audio-driven co-speech gesture generation with diffusion models,” in Proceedings of the 32nd International Joint Conference on Artificial Intelligence, IJCAI 2023 . ijcai.org, 2023
2023
Closest in time.
Z. Allen-Zhu and Y. Li, “Towards understanding ensemble, knowledge distillation and self-distillation in deep learning,” in International Conference on Learning Representations , 2023
2023
Closest in time.
Q. Zhao, C. Zheng, M. Liu, P. Wang, and C. Chen, “Poseformerv2: Exploring frequency domain for efficient and robust 3d human pose estimation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023
2023
Closest in time.
X. Qi, C. Liu, M. Sun, L. Li, C. Fan, and X. Yu, “Diverse 3d hand gesture prediction from body dynamics by bilateral hand disentanglement,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023
2023
Closest in time.
A. Padmakumar, J. Thomason, A. Shrivastava, P. Lange, A. Narayan-Chen, S. Gella, R. Piramuthu, G. Tur, and D. Hakkani-Tur, “Teach: Task-driven embodied agents that chat,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 2, 2022, pp. 2017–2025
2025
Closest in time.
H. S. Koppula and A. Saxena, “Anticipating human activities for reactive robotic response.” in IROS . Tokyo, 2013, p. 2071
2071
Closest in time.