Fetching the paper…
Reading the bibliography…
In this paper, we study the associations between human faces and voices.
Jones, B., Kabanoff, B.: Eye movements in auditory space perception. Attention, Perception, & Psychophysics 17
1975
Earlier work this paper cites.
McGurk, H., MacDonald, J.: Hearing lips and seeing voices. Nature 264
1976
Earlier work this paper cites.
Shelton, B.R., Searle, C.L.: The influence of vision on the absolute identification of sound-source position. Perception & Psychophysics 28
1980
Earlier work this paper cites.
Gaver, W.W.: What in the world do we hear?: An ecological approach to auditory event perception. Ecological psychology 5
1993
Earlier work this paper cites.
Brookes, H., Slater, A., Quinn, P.C., Lewkowicz, D.J., Hayes, R., Brown, E.: Three-month-old infants learn arbitrary auditory–visual pairings between voices and faces. Infant and Child Development 10
2001
Earlier work this paper cites.
Kamachi, M., Hill, H., Lander, K., Vatikiotis-Bateson, E.: “Putting the face to the voice”: Matching identity across modality. Current Biology 13
2003
Earlier work this paper cites.
Lachs, L., Pisoni, D.B.: Crossmodal source identification in speech perception. Ecological Psychology 16
2004
Earlier work this paper cites.
Chopra, S., Hadsell, R., LeCun, Y.: Learning a similarity metric discriminatively, with application to face verification. In: CVPR. pp. 539–546 (2005)
2005
Earlier work this paper cites.
von Kriegstein, K., Kleinschmidt, A., Sterzer, P., Giraud, A.L.: Interaction of face and voice areas during speaker recognition. J. Cogn. Neurosci. 17
2005
Earlier work this paper cites.
Campanella, S., Belin, P.: Integrating face and voice in person perception. Trends in Cognitive Sciences 11
2007
Earlier work this paper cites.
van der Maaten, L., Hinton, G.: Visualizing data using t-SNE. JMLR 9
2008
Earlier work this paper cites.
Joassin, F., Pesenti, M., Maurage, P., Verreckt, E., Bruyer, R., Campanella, S.: Cross-modal interactions between human faces and voices involved in person recognition. Cortex 47
2011
Earlier work this paper cites.
Ngiam, J., Khosla, A., Kim, M., Nam, J., Lee, H., Ng, A.Y.: Multimodal deep learning. In: ICML. pp. 689–696 (2011)
2011
Earlier work this paper cites.
Sliwa, J., Duhamel, J.R., Pascalis, O., Wirth, S.: Spontaneous voice–face identity matching by rhesus monkeys for familiar conspecifics and humans. PNAS 108
2011
Earlier work this paper cites.
Torralba, A., Efros, A.A.: Unbiased look at dataset bias. In: CVPR. pp. 1521–1528 (2011)
2011
Earlier work this paper cites.
Mavica, L.W., Barenholtz, E.: Matching voice and face identity from static images. J. Exp. Psychol. Hum. Percept. Perform. 39
2013
Cited alongside, same era.
2014
Cited alongside, same era.
2014
Cited alongside, same era.
Doersch, C., Gupta, A., Efros, A.A.: Unsupervised visual representation learning by context prediction. In: ICCV. pp. 1422–1430 (2015)
2015
Cited alongside, same era.
Gebru, I.D., Ba, S., Evangelidis, G., Horaud, R.: Tracking the active speaker based on a joint audio-visual observation model. In: ICCV Workshop. pp. 15–21 (2015)
Solèr, M., Bazin, J.C., Wang, O., Krause, A., Sorkine-Hornung, A.: Suggesting sounds for images from video collections. In: ECCV Workshop. pp. 900–917 (2016)
2016
Later among the works it cites.
Arandjelovic, R., Zisserman, A.: Look, listen and learn. In: ICCV. pp. 609–617 (2017)
2017
Later among the works it cites.
Bau, D., Zhou, B., Khosla, A., Oliva, A., Torralba, A.: Network dissection: Quantifying interpretability of deep visual representations. In: CVPR. pp. 3319–3327 (2017)
2017
Later among the works it cites.
Chen, W., Chen, X., Zhang, J., Huang, K.: Beyond triplet loss: A deep quadruplet network for person re-identification. In: CVPR. pp. 1320–1329 (2017)
2017
Later among the works it cites.
Chung, J.S., Senior, A.W., Vinyals, O., Zisserman, A.: Lip reading sentences in the wild. In: CVPR. pp. 3444–3453 (2017)
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
Hoffer, E., Ailon, N.: Deep metric learning using triplet network. In: SIMBAD. pp. 84–92 (2015)
2015
Cited alongside, same era.
Liu, Z., Luo, P., Wang, X., Tang, X.: Deep learning face attributes in the wild. In: ICCV. pp. 3730–3738 (2015)
2015
Cited alongside, same era.
Parkhi, O.M., Vedaldi, A., Zisserman, A.: Deep face recognition. In: BMVC. pp. 41.1–41.12 (2015)
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2015
Cited alongside, same era.
Aytar, Y., Vondrick, C., Torralba, A.: Soundnet: Learning sound representations from unlabeled video. In: NIPS. pp. 892–900 (2016)
2016
Cited alongside, same era.
Owens, A., Isola, P., McDermott, J.H., Torralba, A., Adelson, E.H., Freeman, W.T.: Visually indicated sounds. In: CVPR. pp. 2405–2413 (2016)
2016
Cited alongside, same era.
2017
Later among the works it cites.
Karras, T., Aila, T., Laine, S., Herva, A., Lehtinen, J.: Audio-driven facial animation by joint end-to-end learning of pose and emotion. ACM Trans. Graph. 36
2017
Later among the works it cites.
Nagrani, A., Chung, J.S., Zisserman, A.: Voxceleb: A large-scale speaker identification dataset. In: INTERSPEECH. pp. 2616–2620 (2017)
2017
Later among the works it cites.
Suwajanakorn, S., Seitz, S.M., Kemelmacher-Shlizerman, I.: Synthesizing obama: learning lip sync from audio. ACM Trans. Graph. 36
2017
Later among the works it cites.
Taylor, S.L., Kim, T., Yue, Y., Mahler, M., Krahe, J., Rodriguez, A.G., Hodgins, J.K., Matthews, I.A.: A deep learning approach for generalized speech animation. ACM Trans. Graph. 36
2017
Later among the works it cites.
Tzeng, E., Hoffman, J., Saenko, K., Darrell, T.: Adversarial discriminative domain adaptation. In: CVPR. pp. 2962–2971 (2017)
2017
Later among the works it cites.
Nagrani, A., Albanie, S., Zisserman, A.: Seeing voices and hearing faces: Cross-modal biometric matching. In: CVPR. pp. 8427–8436 (2018)
2018
Closest in time.
2018
Closest in time.
Wu, Z., Singh, B., Davis, L.S., Subrahmanian, V.S.: Deception detection in videos. In: AAAI (2018)
2018
Closest in time.