Fetching the paper…
Reading the bibliography…
This work seeks the possibility of generating the human face from voice solely based on the audio-visual data without any human-labeled annotations.
Determination of the vocal-tract shape from measured formant frequencies
Paul Mermelstein · 1967
Earlier work this paper cites.
Understanding face recognition
Vicki Bruce and Andy Young · 1986
Earlier work this paper cites.
Evidence for nonlinear sound production mechanisms in the vocal tract
HM Teager and SM Teager · 1990
Earlier work this paper cites.
Dlib-ml: A machine learning toolkit
Davis E. King · 2009
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Conditional generative adversarial nets
Mehdi Mirza and Simon Osindero · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Deep face recognition
Omkar M Parkhi, Andrea Vedaldi, Andrew Zisserman, et al · 2015
Earlier work this paper cites.
“hearing faces and seeing voices”: Amodal coding of person identity in the human brain
Bashar Awwad Shiekh Hasan, Mitchell Valdes-Sosa, Joachim Gross, and Pascal Belin · 2016
Earlier work this paper cites.
Generative adversarial text to image synthesis
Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee · 2016
Earlier work this paper cites.
Lip reading sentences in the wild
Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman · 2017
Earlier work this paper cites.
Synthesizing normalized faces from facial identity features
Forrester Cole, David Belanger, Dilip Krishnan, Aaron Sarna, Inbar Mosseri, and William T Freeman · 2017
Cited alongside, same era.
Arbitrary style transfer in real-time with adaptive instance normalization
Xun Huang and Serge Belongie · 2017
Cited alongside, same era.
Voxceleb: a large-scale speaker identification dataset
Joon Son Chung Nagrani, Arsha and Andrew Zisserman · 2017
Cited alongside, same era.
Conditional image synthesis with auxiliary classifier gans
Augustus Odena, Christopher Olah, and Jonathon Shlens · 2017
Cited alongside, same era.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Joy Buolamwini and Timnit Gebru · 2018
Cited alongside, same era.
Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation
Self-attention generative adversarial networks
Han Zhang, Ian Goodfellow, Dimitris Metaxas, and Augustus Odena · 2018
Later among the works it cites.
Uncovering and mitigating algorithmic bias through learned latent structure
Alexander Amini, Ava Soleimany, Wilko Schwarting, Sangeeta Bhatia, and Daniela Rus · 2019
Later among the works it cites.
Large scale gan training for high fidelity natural image synthesis
Andrew Brock, Jeff Donahue, and Karen Simonyan · 2019
Later among the works it cites.
Wav2pix: speech-conditioned face generation using generative adversarial networks
Amanda Duarte, Francisco Roldan, Miquel Tubau, Janna Escur, Santiago Pascual, Amaia Salvador, Eva Mohedano, Kevin McGuinness, Jordi Torres, and Xavier Giro-i Nieto · 2019
Later among the works it cites.
The relativistic discriminator: a key element missing from standard gan
Alexia Jolicoeur-Martineau · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T Freeman, and Michael Rubinstein · 2018
Cited alongside, same era.
Which training methods for gans do actually converge?
Lars Mescheder, Andreas Geiger, and Sebastian Nowozin · 2018
Cited alongside, same era.
cGANs with projection discriminator
Takeru Miyato and Masanori Koyama · 2018
Cited alongside, same era.
Spectral normalization for generative adversarial networks
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida · 2018
Cited alongside, same era.
Speaker recognition from raw waveform with sincnet
Mirco Ravanelli and Yoshua Bengio · 2018
Cited alongside, same era.
Deep audio-visual speech recognition
Triantafyllos Afouras, Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman
Cited in the paper.
Deep lip reading: a comparison of models and an online application
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman
Cited in the paper.
Tero Karras, Samuli Laine, and Timo Aila · 2019
Later among the works it cites.
Speech2face: Learning the face behind a voice
Tae-Hyun Oh, Tali Dekel, Changil Kim, Inbar Mosseri, William T Freeman, Michael Rubinstein, and Wojciech Matusik · 2019
Later among the works it cites.
Learning problem-agnostic speech representations from multiple self-supervised tasks
Santiago Pascual, Mirco Ravanelli, Joan Serrà, Antonio Bonafonte, and Yoshua Bengio · 2019
Later among the works it cites.
Noise-tolerant audio-visual online person verification using an attention-based neural network fusion
Suwon Shon, Tae-Hyun Oh, and James Glass · 2019
Later among the works it cites.
Faces and voices in the brain: a modality-general person-identity representation in superior temporal sulcus
Maria Tsantani, Nikolaus Kriegeskorte, Carolyn McGettigan, and Lúcia Garrido · 2019
Later among the works it cites.