Fetching the paper…
Reading the bibliography…
We present TANGO, a framework for generating co-speech body-gesture videos.
Optical flow estimation
David Fleet and Yair Weiss · 2006
Earlier work this paper cites.
Motion graphs
Lucas Kovar, Michael Gleicher, and Frédéric H. Pighin · 2008
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Flownet 2.0: Evolution of optical flow estimation with deep networks
Eddy Ilg, Nikolaus Mayer, Tonmoy Saikia, Margret Keuper, Alexey Dosovitskiy, and Thomas Brox · 2017
Earlier work this paper cites.
Video frame synthesis using deep voxel flow
Ziwei Liu, Raymond A. Yeh, Xiaoou Tang, Yiming Liu, and Aseem Agarwala · 2017
Earlier work this paper cites.
Super slomo: High quality estimation of multiple intermediate frames for video interpolation
Huaizu Jiang, Deqing Sun, Varun Jampani, Ming-Hsuan Yang, Erik G. Learned-Miller, and Jan Kautz · 2018
Earlier work this paper cites.
Context-aware synthesis for video frame interpolation
Simon Niklaus and Feng Liu · 2018
Earlier work this paper cites.
Everybody dance now
Caroline Chan, Shiry Ginosar, Tinghui Zhou, and Alexei A Efros · 2019
Earlier work this paper cites.
Learning individual styles of conversational gesture
Shiry Ginosar, Amir Bar, Gefen Kohavi, Caroline Chan, Andrew Owens, and Jitendra Malik · 2019
Earlier work this paper cites.
Mediapipe: A framework for building perception pipelines
Camillo Lugaresi, Jiuqiang Tang, Hadon Nash, Chris McClanahan, Esha Uboweja, Michael Hays, Fan Zhang, Chuo-Ling Chang, Ming Guang Yong, Juhyun Lee, et al · 2019
Earlier work this paper cites.
Expressive body capture: 3D hands, face, and body from a single image
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black · 2019
Earlier work this paper cites.
Video enhancement with task-oriented flow
Tianfan Xue, Baian Chen, Jiajun Wu, Donglai Wei, and William T. Freeman · 2019
Earlier work this paper cites.
On the continuity of rotation representations in neural networks
Yi Zhou, Connelly Barnes, Jingwan Lu, Jimei Yang, and Hao Li · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Earlier work this paper cites.
Improved baselines with momentum contrastive learning
Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He · 2020
Earlier work this paper cites.
Mmsegmentation: Openmmlab semantic segmentation toolbox and benchmark, 2020
MMSegmentation Contributors · 2020
Earlier work this paper cites.
Softmax splatting for video frame interpolation
Simon Niklaus and Feng Liu · 2020
Cited alongside, same era.
BMBC: bilateral motion estimation with bilateral cost volume for video interpolation
Junheum Park, Keunsoo Ko, Chul Lee, and Chang-Su Kim · 2020
Cited alongside, same era.
A lip sync expert is all you need for speech to lip generation in the wild
KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar · 2020
Cited alongside, same era.
Speech gesture generation from the trimodal context of text, audio, and speaker identity
Youngwoo Yoon, Bok Cha, Joo-Haeng Lee, Minsu Jang, Jaeyeon Lee, Jaehong Kim, and Geehyuk Lee · 2020
Cited alongside, same era.
Audio2gestures: Generating diverse gestures from speech audio with conditional variational autoencoders
Jing Li, Di Kang, Wenjie Pei, Xuefei Zhe, Ying Zhang, Zhenyu He, and Linchao Bao · 2021
Cited alongside, same era.
FILM: frame interpolation for large motion
Fitsum Reda, Janne Kontkanen, Eric Tabellion, Deqing Sun, Caroline Pantofaru, and Brian Curless · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Later among the works it cites.
Motionclip: Exposing human motion generation to clip space
Guy Tevet, Brian Gordon, Amir Hertz, Amit H Bermano, and Daniel Cohen-Or · 2022
Later among the works it cites.
MCVD - masked conditional video diffusion for prediction, generation, and interpolation
Vikram Voleti, Alexia Jolicoeur-Martineau, and Chris Pal · 2022
Later among the works it cites.
Optimizing video prediction via video frame interpolation
Yue Wu, Qiang Wen, and Qifeng Chen · 2022
Later among the works it cites.
Audio-driven neural gesture reenactment with video motion graphs
Yang Zhou, Jimei Yang, Dingzeyu Li, Jun Saito, Deepali Aneja, and Evangelos Kalogerakis · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Junheum Park, Chul Lee, and Chang-Su Kim · 2021
Cited alongside, same era.
Speech drives templates: Co-speech gesture synthesis with learned templates
Shenhan Qian, Zhi Tu, Yihao Zhi, Wen Liu, and Shenghua Gao · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
XVFI: extreme video frame interpolation
Hyeonjun Sim, Jihyong Oh, and Munchurl Kim · 2021
Cited alongside, same era.
St-mfnet: A spatio-temporal multi-flow network for frame interpolation
Duolikun Danier, Fan Zhang, and David R. Bull · 2022
Cited alongside, same era.
Tm2t: Stochastic and tokenized modeling for the reciprocal generation of 3d human motions and texts
Chuan Guo, Xinxin Zuo, Sen Wang, and Li Cheng · 2022
Cited alongside, same era.
Nemf: Neural motion fields for kinematic animation
Chengan He, Jun Saito, James Zachary, Holly Rushmeier, and Yi Zhou · 2022
Cited alongside, same era.
Later among the works it cites.
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Yuwei Guo, Ceyuan Yang, Anyi Rao, Yaohui Wang, Yu Qiao, Dahua Lin, and Bo Dai · 2023
Later among the works it cites.
Animate anyone: Consistent and controllable image-to-video synthesis for character animation
Li Hu, Xin Gao, Peng Zhang, Ke Sun, Bang Zhang, and Liefeng Bo · 2023
Later among the works it cites.
A comprehensive review of data-driven co-speech gesture generation
Simbarashe Nyatsanga, Taras Kucherenko, Chaitanya Ahuja, Gustav Eje Henter, and Michael Neff · 2023
Later among the works it cites.
Bodyformer: Semantics-guided 3d body gesture synthesis with transformer
Kunkun Pang, Dafei Qin, Yingruo Fan, Julian Habekost, Takaaki Shiratori, Junichi Yamagishi, and Taku Komura · 2023
Later among the works it cites.
Biformer: Learning bilateral motion estimation via bilateral transformer for 4k video frame interpolation
Junheum Park, Jintae Kim, and Chang-Su Kim · 2023
Later among the works it cites.
Generating holistic 3d human motion from speech
Hongwei Yi, Hualin Liang, Yifei Liu, Qiong Cao, Yandong Wen, Timo Bolkart, Dacheng Tao, and Michael J Black · 2023
Later among the works it cites.
Taming diffusion models for audio-driven co-speech gesture generation
Lingting Zhu, Xian Liu, Xuanyu Liu, Rui Qian, Ziwei Liu, and Lequan Yu · 2023
Later among the works it cites.
LDMVFI: video frame interpolation with latent diffusion models
Duolikun Danier, Fan Zhang, and David R. Bull · 2024
Closest in time.
Video interpolation with diffusion models
Siddhant Jain, Daniel Watson, Eric Tabellion, Aleksander Holynski, Ben Poole, and Janne Kontkanen · 2024
Closest in time.