Fetching the paper…
Reading the bibliography…
The thud of a bouncing ball, the onset of speech as lips open -- when visual and audio events occur together, it suggests that there might be a common, underlying event that produced both signals.
Some experiments on the recognition of speech, with one and with two ears
Cherry, E.C.: · 1953
Earlier work this paper cites.
Hearing lips and seeing voices
McGurk, H., MacDonald, J.: · 1976
Earlier work this paper cites.
Learning classification with unlabeled data
de Sa, V.R.: · 1994
Earlier work this paper cites.
Auditory scene analysis: The perceptual organization of sound
Bregman, A.S.: · 1994
Earlier work this paper cites.
Factorial hidden markov models
Ghahramani, Z., Jordan, M.I.: · 1996
Earlier work this paper cites.
Sound alters visual motion perception
Sekuler, R.: · 1997
Earlier work this paper cites.
Lip synchronization of speech
McAllister, D.F., Rodman, R.D., Bitzer, D.L., Freeman, A.S.: · 1997
Earlier work this paper cites.
Is primitive av coherence an aid to segment the scene?
Barker, J.P., Berthommier, F., Schwartz, J.L.: · 1998
Earlier work this paper cites.
Audio vision: Using audio-visual synchrony to locate sounds
Hershey, J.R., Movellan, J.R.: · 1999
Earlier work this paper cites.
Learning joint statistical models for audio-visual fusion and segregation
Fisher III, J.W., Darrell, T., Freeman, W.T., Viola, P.A.: · 2000
Earlier work this paper cites.
Audio-visual segmentation and “the cocktail party effect”
Darrell, T., Fisher, J.W., Viola, P.: · 2000
Earlier work this paper cites.
Sensory modalities are not separate modalities: plasticity and interactions
Shimojo, S., Shams, L.: · 2001
Earlier work this paper cites.
One microphone source separation
Roweis, S.T.: · 2001
Earlier work this paper cites.
Rapid object detection using a boosted cascade of simple features
Viola, P., Jones, M.: · 2001
Earlier work this paper cites.
Audio-visual scene analysis: evidence for a” very-early” integration process in audio-visual speech perception
Schwartz, J.L., Berthommier, F., Savariaux, C.: · 2002
Earlier work this paper cites.
Tsp speech database
Kabal, P.: · 2002
Earlier work this paper cites.
Audio-visual graphical models for speech processing
Hershey, J., Attias, H., Jojic, N., Kristjansson, T.: · 2004
Earlier work this paper cites.
The development of embodied cognition: Six lessons from babies
Smith, L., Gasser, M.: · 2005
Earlier work this paper cites.
Pixels that sound
Kidron, E., Schechner, Y.Y., Elad, M.: · 2005
Earlier work this paper cites.
Performance measurement in blind audio source separation
Vincent, E., Gribonval, R., Févotte, C.: · 2006
Earlier work this paper cites.
An audio-visual corpus for speech perception and automatic speech recognition
Cooke, M., Barker, J., Cunningham, S., Shao, X.: · 2006
Earlier work this paper cites.
Harmony in motion
Barzelay, Z., Schechner, Y.Y.: · 2007
Earlier work this paper cites.
Monaural sound source separation by nonnegative matrix factorization with temporal continuity and sparseness criteria
Virtanen, T.: · 2007
Earlier work this paper cites.
Fusion and combination in audio-visual integration
Omata, K., Mogi, K.: · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: · 2009
Earlier work this paper cites.
Is seeing believing? (2010)
British Broadcasting Corporation: · 2010
Earlier work this paper cites.
Monaural speech separation and recognition challenge
Cooke, M., Hershey, J.R., Rennie, S.J.: · 2010
Earlier work this paper cites.
Blind audiovisual source separation based on sparse redundant representations
Casanovas, A.L., et al.: · 2010
Cited alongside, same era.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M., Hyvärinen, A.: · 2010
Cited alongside, same era.
Multimodal deep learning
Ngiam, J., Khosla, A., Kim, M., Nam, J., Lee, H., Ng, A.Y.: · 2011
Cited alongside, same era.
Binding and unbinding the auditory and visual streams in the mcgurk effect
Nahorna, O., Berthommier, F., Schwartz, J.L.: · 2012
Cited alongside, same era.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Soomro, K., Zamir, A.R., Shah, M.: · 2012
Cited alongside, same era.
Speaker separation using visually-derived binary masks
Learning deep features for discriminative localization
Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., Torralba, A.: · 2016
Later among the works it cites.
Pose from action: Unsupervised learning of pose features based on motion
Purushwalkam, S., Gupta, A.: · 2016
Later among the works it cites.
Image-to-image translation with conditional adversarial networks
Isola, P., Zhu, J.Y., Zhou, T., Efros, A.A.: · 2016
Later among the works it cites.
Look, listen and learn
Arandjelović, R., Zisserman, A.: · 2017
Later among the works it cites.
Self-supervised video representation learning with odd-one-out networks
Fernando, B., Bilen, H., Gavves, E., Gould, S.: · 2017
Later among the works it cites.
Lip reading sentences in the wild
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Khan, F., Milner, B.: · 2013
Cited alongside, same era.
Lin, M., Chen, Q., Yan, S.: · 2013
Cited alongside, same era.
Audiovisual speech source separation: An overview of key methodologies
Rivet, B., et al.: · 2014
Cited alongside, same era.
Two-stream convolutional networks for action recognition in videos
Simonyan, K., Zisserman, A.: · 2014
Cited alongside, same era.
Large-scale video classification with convolutional neural networks
Karpathy, A., Toderici, G., Shetty, S., Leung, T., Sukthankar, R., Fei-Fei, L.: · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, K., Zisserman, A.: · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D.P., Ba, J.: · 2014
Cited alongside, same era.
Chung, J.S., Senior, A., Vinyals, O., Zisserman, A.: · 2017
Later among the works it cites.
Deep attractor network for single-microphone speaker separation
Chen, Z., Luo, Y., Mesgarani, N.: · 2017
Later among the works it cites.
Permutation invariant training of deep models for speaker-independent multi-talker speech separation
Yu, D., Kolbæk, M., Tan, Z.H., Jensen, J.: · 2017
Later among the works it cites.
Audio-visual object localization and separation using low-rank and sparsity
Pu, J., et al.: · 2017
Later among the works it cites.
Audio-visual speech enhancement using multimodal deep convolutional neural networks
Hou, J.C., Wang, S.S., Lai, Y.H., Tsao, Y., Chang, H.W., Wang, H.M.: · 2017
Later among the works it cites.
Seeing through noise: Speaker separation and enhancement using visually-derived speech
Gabbay, A., Ephrat, A., Halperin, T., Peleg, S.: · 2017
Later among the works it cites.
Visual speech enhancement using noise-invariant training
Gabbay, A., Shamir, A., Peleg, S.: · 2017
Later among the works it cites.
Arandjelović, R., Zisserman, A.: · 2017
Later among the works it cites.
Learning sight from sound: Ambient sound provides supervision for visual learning
Owens, A., Wu, J., McDermott, J.H., Freeman, W.T., Torralba, A.: · 2017
Later among the works it cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J., Zisserman, A.: · 2017
Later among the works it cites.
Audio set: An ontology and human-labeled dartaset for audio events
Gemmeke, J.F., Ellis, D.P., Freedman, D., Jansen, A., Lawrence, W., Moore, R.C., Plakal, M., Ritter, M.: · 2017
Later among the works it cites.
The kinetics human action video dataset
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., et al.: · 2017
Later among the works it cites.
Michelsanti, D., Tan, Z.H.: · 2017
Later among the works it cites.
Voxceleb: a large-scale speaker identification dataset
Nagrani, A., Chung, J.S., Zisserman, A.: · 2017
Later among the works it cites.
Learning and using the arrow of time
Wei, D., Lim, J.J., Zisserman, A., Freeman, W.T.: · 2018
Closest in time.
Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation
Ephrat, A., Mosseri, I., Lang, O., Dekel, T., Wilson, K., Hassidim, A., Freeman, W.T., Rubinstein, M.: · 2018
Closest in time.
The conversation: Deep audio-visual speech enhancement
Afouras, T., Chung, J.S., Zisserman, A.: · 2018
Closest in time.
Zhao, H., Gan, C., Rouditchenko, A., Vondrick, C., McDermott, J., Torralba, A.: · 2018
Closest in time.
Learning to Separate Object Sounds by Watching Unlabeled Video
Gao, R., Feris, R., Grauman, K.: · 2018
Closest in time.
Learning to localize sound source in visual scenes
Senocak, A., Oh, T.H., Kim, J., Yang, M.H., Kweon, I.S.: · 2018
Closest in time.
Voxceleb2: Deep speaker recognition
Chung, J.S., Nagrani, A., Zisserman, A.: · 2018
Closest in time.