Fetching the paper…
Reading the bibliography…
Talking face generation aims to synthesize a face video with precise lip synchronization as well as a smooth transition of facial motion over the entire video via the given speech clip and facial image.
Asymptotic evaluation of certain markov process expectations for large time. iv
Monroe D Donsker and SR Srinivasa Varadhan · 1983
Earlier work this paper cites.
Estimating mutual information
Alexander Kraskov, Harald Stögbauer, and Peter Grassberger · 2004
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli · 2004
Earlier work this paper cites.
Edgeworth approximation of multivariate differential entropy
Marc M Van Hulle · 2005
Earlier work this paper cites.
An audio-visual corpus for speech perception and automatic speech recognition
Martin Cooke, Jon Barker, Stuart Cunningham, and Xu Shao · 2006
Earlier work this paper cites.
Robust discriminant analysis based on nonparametric maximum entropy
Ran He, Bao-Gang Hu, and Xiao-Tong Yuan · 2009
Earlier work this paper cites.
Feature selection with dynamic mutual information
Huawen Liu, Jigui Sun, Lei Liu, and Huijie Zhang · 2009
Earlier work this paper cites.
Mib: Using mutual information for biclustering gene expression data
Neelima Gupta and Seema Aggarwal · 2010
Earlier work this paper cites.
Supervised feature selection by clustering using conditional mutual information-based distances
José Martínez Sotoca and Filiberto Pla · 2010
Earlier work this paper cites.
Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion
Pascal Vincent, Hugo Larochelle, Isabelle Lajoie, Yoshua Bengio, and Pierre-Antoine Manzagol · 2010
Cited alongside, same era.
A compact local binary pattern using maximization of mutual information for face analysis
Bongjin Jun, Taewan Kim, and Daijin Kim · 2011
Cited alongside, same era.
Tighter variational representations of f-divergences via restriction to probability measures
Avraham Ruderman, Mark Reid, Darío García-García, and James Petterson · 2012
Cited alongside, same era.
Robust recognition via information theoretic learning
Ran He, Baogang Hu, Xiaotong Yuan, and Liang Wang · 2014
Cited alongside, same era.
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel · 2016
Cited alongside, same era.
Obamanet: Photo-realistic lip-sync from text
Rithesh Kumar, Jose Sotelo, Kundan Kumar, Alexandre de Brébisson, and Yoshua Bengio · 2017
Later among the works it cites.
Synthesizing obama: learning lip sync from audio
Supasorn Suwajanakorn, Steven M Seitz, and Ira Kemelmacher-Shlizerman · 2017
Later among the works it cites.
Mutual information neural estimation
Mohamed Ishmael Belghazi, Aristide Baratin, Sai Rajeshwar, Sherjil Ozair, Yoshua Bengio, R. Devon Hjelm, and Aaron C. Courville · 2018
Closest in time.
Lip movements generation at a glance
Lele Chen, Zhiheng Li, Ross K Maddox, Zhiyao Duan, and Chenliang Xu · 2018
Closest in time.
Learning deep representations by mutual information estimation and maximization
R Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Karan Grewal, Adam Trischler, and Yoshua Bengio · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lip reading in the wild
Joon Son Chung and Andrew Zisserman · 2016
Cited alongside, same era.
Perceptual losses for real-time style transfer and super-resolution
Justin Johnson, Alexandre Alahi, and Li Fei-Fei · 2016
Cited alongside, same era.
You said that?
Joon Son Chung, Amir Jamaludin, and Andrew Zisserman · 2017
Cited alongside, same era.
End-to-end speech-driven facial animation with temporal gans
Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic · 2018
Closest in time.
Hierarchical cross-modal talking face generation with dynamic pixel-wise loss
Lele Chen, Ross K Maddox, Zhiyao Duan, and Chenliang Xu · 2019
Closest in time.
Talking face generation by adversarially disentangled audio-visual representation
Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang · 2019
Closest in time.