Fetching the paper…
Reading the bibliography…
We present ObamaNet, the first architecture that generates both audio and synchronized photo-realistic lip-sync videos from any new text.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros · 2016
Earlier work this paper cites.
Face2face: Real-time face capture and reenactment of rgb videos
J. Thies, M. Zollhöfer, M. Stamminger, C. Theobalt, and M. Nießner · 2016
Cited alongside, same era.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
Tero Karras, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen · 2017
Cited alongside, same era.
Char2wav: End-to-end speech synthesis
Jose Sotelo, Soroush Mehri, Kundan Kumar, Joao Felipe Santos, Kyle Kastner, Aaron Courville, and Yoshua Bengio · 2017
Closest in time.
Synthesizing obama: learning lip sync from audio
Supasorn Suwajanakorn, Steven M Seitz, and Ira Kemelmacher-Shlizerman · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…