Fetching the paper…
Reading the bibliography…
How much can we infer about a person's looks from the way they speak? In this paper, we study the task of reconstructing a facial image of a person from a short audio recording of that person speaking.
The speech chain
P. B. Denes, P. Denes, and E. Pinson · 1993
Earlier work this paper cites.
Minimizing disagreement for self-supervised classification
V. R. de Sa · 1994
Earlier work this paper cites.
Putting the face to the voice: Matching identity across modality
M. Kamachi, H. Hill, K. Lander, and E. Vatikiotis-Bateson · 2003
Earlier work this paper cites.
Visualizing data using T-SNE
L. van der Maaten and G. Hinton · 2008
Earlier work this paper cites.
Dlib-ml: A machine learning toolkit
D. E. King · 2009
Earlier work this paper cites.
Automatic speaker age and gender recognition in the car for tailoring dialog and mobile services
M. Feld, F. Burkhardt, and C. Müller · 2010
Earlier work this paper cites.
Speaker age estimation and gender detection based on supervised non-negative matrix factorization
M. H. Bahari and H. Van Hamme · 2011
Earlier work this paper cites.
Multimodal deep learning
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng · 2011
Earlier work this paper cites.
Deep canonical correlation analysis
G. Andrew, R. Arora, J. Bilmes, and K. Livescu · 2013
Earlier work this paper cites.
Tracking the active speaker based on a joint audio-visual observation model
I. D. Gebru, S. Ba, G. Evangelidis, and R. Horaud · 2015
Earlier work this paper cites.
Speaker height estimation from speech: Fusing spectral regression and statistical acoustic models
J. H. Hansen, K. Williams, and H. Bořil · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
The Discovery of perceptual structure from visual co-occurrences in space and time
P. J. Isola · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Deep face recognition
O. M. Parkhi, A. Vedaldi, and A. Zisserman · 2015
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Y. Aytar, C. Vondrick, and A. Torralba · 2016
Earlier work this paper cites.
Learning aligned cross-modal representations from weakly aligned data
L. Castrejon, Y. Aytar, C. Vondrick, H. Pirsiavash, and A. Torralba · 2016
Cited alongside, same era.
Coupled generative adversarial networks
M.-Y. Liu and O. Tuzel · 2016
Cited alongside, same era.
Visually indicated sounds
A. Owens, P. Isola, J. H. McDermott, A. Torralba, E. H. Adelson, and W. T. Freeman · 2016
Cited alongside, same era.
Matching novel face and voice identity using static and dynamic facial images
H. M. Smith, A. K. Dunn, T. Baguley, and P. C. Stacey · 2016
Cited alongside, same era.
Suggesting sounds for images from video collections
M. Solèr, J. C. Bazin, O. Wang, A. Krause, and A. Sorkine-Hornung · 2016
Cited alongside, same era.
Attribute2image: Conditional image generation from visual attributes
X. Yan, J. Yang, K. Sohn, and H. Lee · 2016
Cited alongside, same era.
Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation
A. Ephrat, I. Mosseri, O. Lang, T. Dekel, K. Wilson, A. Hassidim, W. T. Freeman, and M. Rubinstein · 2018
Later among the works it cites.
Face-voice matching using cross-modal embeddings
S. Horiguchi, N. Kanda, and K. Nagamatsu · 2018
Later among the works it cites.
On learning associations of faces and voices
C. Kim, H. V. Shin, T.-H. Oh, A. Kaspar, M. Elgharib, and W. Matusik · 2018
Later among the works it cites.
Learnable PINs: Cross-modal embeddings for person identity
A. Nagrani, S. Albanie, and A. Zisserman · 2018
Later among the works it cites.
Seeing voices and hearing faces: Cross-modal biometric matching
A. Nagrani, S. Albanie, and A. Zisserman · 2018
Later among the works it cites.
Audio-visual scene analysis with self-supervised multisensory features
A. Owens and A. A. Efros · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Look, listen and learn
R. Arandjelovic and A. Zisserman · 2017
Cited alongside, same era.
Lip reading sentences in the wild
J. S. Chung, A. W. Senior, O. Vinyals, and A. Zisserman · 2017
Cited alongside, same era.
Synthesizing normalized faces from facial identity features
F. Cole, D. Belanger, D. Krishnan, A. Sarna, I. Mosseri, and W. T. Freeman · 2017
Cited alongside, same era.
Putting a face to the voice: Fusing audio and visual signals across a video to determine speakers
K. Hoover, S. Chaudhuri, C. Pantofaru, M. Slaney, and I. Sturdy · 2017
Cited alongside, same era.
Unsupervised learning of disentangled and interpretable representations from sequential data
W.-N. Hsu, Y. Zhang, and J. Glass · 2017
Cited alongside, same era.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
T. Karras, T. Aila, S. Laine, A. Herva, and J. Lehtinen · 2017
Cited alongside, same era.
Later among the works it cites.
Learning sight from sound: Ambient sound provides supervision for visual learning
A. Owens, J. Wu, J. H. McDermott, W. T. Freeman, and A. Torralba · 2018
Later among the works it cites.
Speech-driven expressive talking lips with conditional sequential generative adversarial networks
N. Sadoughi and C. Busso · 2018
Later among the works it cites.
Learning to localize sound source in visual scenes
A. Senocak, T.-H. Oh, J. Kim, M.-H. Yang, and I. S. Kweon · 2018
Later among the works it cites.
Audio to body dynamics
E. Shlizerman, L. Dery, H. Schoen, and I. Kemelmacher-Shlizerman · 2018
Later among the works it cites.
S. Shon, T.-H. Oh, and J. Glass · 2018
Later among the works it cites.
X2Face: A network for controlling face generation using images, audio, and pose codes
O. Wiles, A. S. Koepke, and A. Zisserman · 2018
Later among the works it cites.
Age estimation in short speech utterances based on LSTM recurrent neural networks
R. Zazo, P. S. Nidadavolu, N. Chen, J. Gonzalez-Rodriguez, and N. Dehak · 2018
Later among the works it cites.
The sound of pixels
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. H. McDermott, and A. Torralba · 2018
Later among the works it cites.
Wav2Pix: speech-conditioned face generation using generative adversarial networks
A. Duarte, F. Roldan, M. Tubau, J. Escur, S. Pascual, A. Salvador, E. Mohedano, K. McGuinness, J. Torres, and X. Giro-i-Nieto · 2019
Closest in time.
M. Merler, N. Ratha, R. S. Feris, and J. R. Smith · 2019
Closest in time.
Disjoint mapping network for cross-modal matching of voices and faces
Y. Wen, M. A. Ismail, W. Liu, B. Raj, and R. Singh · 2019
Closest in time.