Fetching the paper…
Reading the bibliography…
With the advance of deep learning technology, automatic video generation from audio or text has become an emerging and promising research topic.
On the Acoustic Structure of Diphthongal Syllables
Hongmo Ren, · 1986
Earlier work this paper cites.
“Timit acoustic-phonetic continuous speech corpus,”
J. Garofolo, Lori Lamel, W. Fisher, Jonathan Fiscus, D. Pallett, N. Dahlgren, and V. Zue, · 1992
Earlier work this paper cites.
“Visual speech synthesis by morphing visemes,”
Tony Ezzat and Tomaso Poggio, · 2000
Earlier work this paper cites.
“Speaker identification on the scotus corpus,”
J. Yuan and M. Liberman, · 2008
Earlier work this paper cites.
“Multi-region probabilistic histograms for robust and scalable identity inference,”
Conrad Sanderson and Brian C Lovell, · 2009
Earlier work this paper cites.
“Dynamic units of visual speech,”
Sarah L Taylor, Moshe Mahler, Barry-John Theobald, and Iain Matthews, · 2012
Earlier work this paper cites.
“Synthesizing obama: learning lip sync from audio,”
Supasorn Suwajanakorn, Steven M Seitz, and Ira Kemelmacher-Shlizerman, · 2017
Cited alongside, same era.
“A deep learning approach for generalized speech animation,”
Sarah Taylor, Taehwan Kim, Yisong Yue, Moshe Mahler, James Krahe, Anastasio Garcia Rodriguez, Jessica Hodgins, and Iain Matthews, · 2017
Cited alongside, same era.
“OpenPose: realtime multi-person 2D pose estimation using Part Affinity Fields,”
Zhe Cao, Gines Hidalgo, Tomas Simon, Shih-En Wei, and Yaser Sheikh, · 2018
Cited alongside, same era.
“Video-to-video synthesis,”
Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Guilin Liu, Andrew Tao, Jan Kautz, and Bryan Catanzaro, · 2018
Cited alongside, same era.
“Neural voice puppetry: Audio-driven facial reenactment,”
Justus Thies, Mohamed Elgharib, Ayush Tewari, Christian Theobalt, and Matthias Nießner, · 2019
Cited alongside, same era.
“Hierarchical cross-modal talking face generation with dynamic pixel-wise loss,”
Lele Chen, Ross K Maddox, Zhiyao Duan, and Chenliang Xu, · 2019
Later among the works it cites.
“Talking face generation by adversarially disentangled audio-visual representation,”
Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang, · 2019
Later among the works it cites.
“Learning individual styles of conversational gesture,”
Shiry Ginosar, Amir Bar, Gefen Kohavi, Caroline Chan, Andrew Owens, and Jitendra Malik, · 2019
Later among the works it cites.
“Deferred neural rendering: Image synthesis using neural textures,”
Justus Thies, Michael Zollhöfer, and Matthias Nießner, · 2019
Later among the works it cites.
“Speech2video synthesis with 3d skeleton regularization and expressive body poses,”
Miao Liao, Sibo Zhang, Peng Wang, Hao Zhu, Xinxin Zuo, and Ruigang Yang, · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Text-based editing of talking-head video,”
Ohad Fried, Ayush Tewari, Michael Zollhöfer, Adam Finkelstein, Eli Shechtman, Dan B Goldman, Kyle Genova, Zeyu Jin, Christian Theobalt, and Maneesh Agrawala, · 2019
Cited alongside, same era.
Wentao Wang, Yan Wang, Jianqing Sun, Qingsong Liu, Jiaen Liang, and Teng Li, · 2020
Later among the works it cites.