Fetching the paper…
Reading the bibliography…
Audio-visual learning, aimed at exploiting the relationship between audio and visual modalities, has drawn considerable attention since deep learning started to be used successfully.
R. V. Shannon, F.-G. Zeng, V. Kamath, J. Wygonski, and M. Ekelid, “Speech recognition with primarily temporal cues,” Science , pp. 303–304, 1995
1995
Earlier work this paper cites.
R. K. Srihari, “Combining text and image information in content-based retrieval,” in Proceedings., International Conference on Image Processing , 1995, pp. 326–329 vol.1
1995
Earlier work this paper cites.
L. R. Long, L. E. Berman, and G. R. Thoma, “Prototype client/server application for biomedical text/image retrieval on the Internet,” in Storage and Retrieval for Still Image and Video Databases IV , 1996, pp. 362 – 372
1996
Earlier work this paper cites.
T. Darrell, J. W. Fisher, and P. Viola, “Audio-visual segmentation and “the cocktail party effect”,” in Advances in Multimodal Interfaces—ICMI 2000 , 2000, pp. 32–40
2000
Earlier work this paper cites.
J. Hershey and J. Movellan, “Audio-vision: Using audio-visual synchrony to locate sounds,” in Advances in Neural Information Processing Systems 12 , 2000, pp. 813–819
2000
Earlier work this paper cites.
S. Dupont and J. Luettin, “Audio-visual speech modeling for continuous speech recognition,” IEEE transactions on multimedia , pp. 141–151, 2000
2000
Earlier work this paper cites.
M. Brand and A. Hertzmann, “Style machines,” in Proceedings of the 27th Annual Conference on Computer Graphics and Interactive Techniques, SIGGRAPH 2000, New Orleans, LA, USA, July 23-28, 2000 , 2000, pp. 183–192
2000
Earlier work this paper cites.
J. W. Fisher III, T. Darrell, W. T. Freeman, and P. A. Viola, “Learning joint statistical models for audio-visual fusion and segregation,” in Advances in Neural Information Processing Systems 13 , 2001, pp. 772–778
2001
Earlier work this paper cites.
2001
Earlier work this paper cites.
G. Potamianos, C. Neti, G. Gravier, A. Garg, and A. W. Senior, “Recent advances in the automatic recognition of audiovisual speech,” Proceedings of the IEEE , pp. 1306–1326, 2003
2003
Earlier work this paper cites.
H. L. Van Trees, Optimum array processing: Part IV of detection, estimation, and modulation theory . John Wiley & Sons, 2004
2004
Earlier work this paper cites.
A. Graves and J. Schmidhuber, “Framewise phoneme classification with bidirectional lstm and other neural network architectures,” Neural Networks , pp. 602–610, 2005
2005
Earlier work this paper cites.
M. Cooke, J. Barker, S. Cunningham, and X. Shao, “An audio-visual corpus for speech perception and automatic speech recognition,” The Journal of the Acoustical Society of America , pp. 2421–2424, 2006
2006
Earlier work this paper cites.
O. Gillet and G. Richard, “Enst-drums: an extensive audio-visual database for drum signals processing.” in ISMIR , 2006, pp. 156–159
2006
Earlier work this paper cites.
J. M. Wang, D. J. Fleet, and A. Hertzmann, “Multifactor gaussian process models for style-content separation,” in Machine Learning, Proceedings of the Twenty-Fourth International Conference (ICML 2007), Corvallis, Oregon, USA, June 20-24, 2007 , 2007, pp. 975–982
2007
Earlier work this paper cites.
G. W. Taylor and G. E. Hinton, “Factored conditional restricted boltzmann machines for modeling motion style,” in Proceedings of the 26th Annual International Conference on Machine Learning, ICML 2009, Montreal, Quebec, Canada, June 14-18, 2009 , 2009, pp. 1025–1032
2009
Earlier work this paper cites.
Y. Bengio, J. Louradour, R. Collobert, and J. Weston, “Curriculum learning,” in Proceedings of the 26th annual international conference on machine learning , 2009, pp. 41–48
2009
Earlier work this paper cites.
C. Sanderson and B. C. Lovell, “Multi-region probabilistic histograms for robust and scalable identity inference,” in International Conference on Biometrics , 2009, pp. 199–208
2009
Earlier work this paper cites.
G. Zhao, M. Barnard, and M. Pietikainen, “Lipreading with local spatiotemporal descriptors,” IEEE Transactions on Multimedia , pp. 1254–1265, 2009
2009
Earlier work this paper cites.
R. He, W.-S. Zheng, and B.-G. Hu, “Maximum correntropy criterion for robust face recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence , pp. 1561–1576, 2010
2010
Earlier work this paper cites.
N. Rasiwasia, J. Costa Pereira, E. Coviello, G. Doyle, G. R. Lanckriet, R. Levy, and N. Vasconcelos, “A new approach to cross-modal multimedia retrieval,” in Proceedings of the 18th ACM International Conference on Multimedia , 2010, pp. 251–260
2010
Earlier work this paper cites.
J. Tilmanne and T. Dutoit, “Expressive gait synthesis using PCA and gaussian modeling,” in Motion in Games - Third International Conference, MIG 2010, Utrecht, The Netherlands, November 14-16, 2010. Proceedings , 2010, pp. 363–374
2010
Earlier work this paper cites.
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng, “Multimodal deep learning,” in Proceedings of the 28th international conference on machine learning (ICML-11) , 2011, pp. 689–696
2011
Earlier work this paper cites.
Y. Bengio, A. Courville, and P. Vincent, “Representation learning: A review and new perspectives,” IEEE transactions on pattern analysis and machine intelligence , pp. 1798–1828, 2013
2013
Earlier work this paper cites.
H. Izadinia, I. Saleemi, and M. Shah, “Multimodal analysis for identification and segmentation of moving-sounding objects,” IEEE Transactions on Multimedia , pp. 378–390, 2013
2013
Earlier work this paper cites.
A. Samadani, E. Kubica, R. Gorbet, and D. Kulic, “Perception and generation of affective hand movements,” I. J. Social Robotics , pp. 35–51, 2013
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Two-Stream Convolutional Networks for Action Recognition in Videos,” in Advances in Neural Information Processing Systems 27 , 2014, pp. 568–576
2014
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Advances in neural information processing systems , 2014, pp. 2672–2680
2014
Earlier work this paper cites.
A. Davis, M. Rubinstein, N. Wadhwa, G. J. Mysore, F. Durand, and W. T. Freeman, “The visual microphone: passive recovery of sound from video,” 2014
2014
Earlier work this paper cites.
A. Zunino, M. Crocco, S. Martelli, A. Trucco, A. Del Bue, and V. Murino, “Seeing the sound: A new multimodal imaging device for computer vision,” in Proceedings of the IEEE International Conference on Computer Vision Workshops , 2015, pp. 6–14
2015
Earlier work this paper cites.
E. Hoffer and N. Ailon, “Deep metric learning using triplet network,” in International Workshop on Similarity-Based Pattern Recognition , 2015, pp. 84–92
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” CoRR , 2015
2015
Earlier work this paper cites.
H. Ninomiya, N. Kitaoka, S. Tamura, Y. Iribe, and K. Takeda, “Integration of deep bottleneck features for audio-visual speech recognition,” in Sixteenth Annual Conference of the International Speech Communication Association , 2015
2015
Earlier work this paper cites.
T. L. Cornu and B. Milner, “Reconstructing intelligible audio speech from visual speech features,” in Sixteenth Annual Conference of the International Speech Communication Association , 2015
2015
Earlier work this paper cites.
N. Harte and E. Gillen, “Tcd-timit: An audio-visual corpus of continuous speech,” IEEE Transactions on Multimedia , pp. 603–615, 2015
2015
Earlier work this paper cites.
I. Anina, Z. Zhou, G. Zhao, and M. Pietikäinen, “Ouluvs2: A multi-view audiovisual database for non-rigid mouth motion analysis,” in Automatic Face and Gesture Recognition (FG), 2015 11th IEEE International Conference and Workshops on , 2015, pp. 1–5
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
M. Wand, J. Koutník, and J. Schmidhuber, “Lipreading with long short-term memory,” in Acoustics, Speech and Signal Processing (ICASSP), 2016 IEEE International Conference on , 2016, pp. 6115–6119
2016
Earlier work this paper cites.
Y. M. Assael, B. Shillingford, S. Whiteson, and N. de Freitas, “Lipnet: Sentence-level lipreading,” arXiv preprint , 2016
2016
Earlier work this paper cites.
G. Trigeorgis, F. Ringeval, R. Brueckner, E. Marchi, M. A. Nicolaou, B. Schuller, and S. Zafeiriou, “Adieu features? end-to-end speech emotion recognition using a deep convolutional recurrent network,” in Acoustics, Speech and Signal Processing (ICASSP), 2016 IEEE International Conference on , 2016, pp. 5200–5204
2016
Earlier work this paper cites.
S. Petridis and M. Pantic, “Prediction-based audiovisual fusion for classification of non-linguistic vocalisations,” IEEE Transactions on Affective Computing , pp. 45–58, 2016
2016
Earlier work this paper cites.
D. Hu, X. Li et al. , “Temporal multimodal learning in audiovisual speech recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 3574–3582
2016
Earlier work this paper cites.
M. Nussbaum-Thom, J. Cui, B. Ramabhadran, and V. Goel, “Acoustic modeling using bidirectional gated recurrent convolutional units,” 2016, pp. 390–394
2016
Earlier work this paper cites.
A. Owens, P. Isola, J. McDermott, A. Torralba, E. H. Adelson, and W. T. Freeman, “Visually indicated sounds,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 2405–2413
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
L. Crnkovic-Friis and L. Crnkovic-Friis, “Generative choreography using deep learning,” CoRR , 2016
2016
Earlier work this paper cites.
D. Holden, J. Saito, and T. Komura, “A deep learning framework for character motion synthesis and editing,” ACM Transactions on Graphics (TOG) , p. 138, 2016
2016
Cited alongside, same era.
Y. Aytar, C. Vondrick, and A. Torralba, “Soundnet: Learning sound representations from unlabeled video,” in Advances in Neural Information Processing Systems , 2016, pp. 892–900
2016
Cited alongside, same era.
J. S. Chung and A. Zisserman, “Lip reading in the wild,” in Asian Conference on Computer Vision , 2016, pp. 87–103
2016
Cited alongside, same era.
2016
Cited alongside, same era.
I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. C. Courville, “Improved training of wasserstein gans,” in NIPS , 2017, pp. 5767–5777
K. Sriskandaraja, V. Sethu, and E. Ambikairajah, “Deep siamese architecture based replay detection for secure voice biometric.” in Interspeech , 2018, pp. 671–675
2018
Later among the works it cites.
X. Wu, R. He, Z. Sun, and T. Tan, “A light cnn for deep face representation with noisy labels,” IEEE Transactions on Information Forensics and Security , pp. 2884–2896, 2018
2018
Later among the works it cites.
A. Nagrani, S. Albanie, and A. Zisserman, “Seeing voices and hearing faces: Cross-modal biometric matching,” CoRR , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
R. Arandjelovic and A. Zisserman, “Look, listen and learn,” in 2017 IEEE International Conference on Computer Vision (ICCV) , 2017, pp. 609–617
2017
Cited alongside, same era.
K. D. Bochen Li, Z. Duan, and G. Sharma, “See and listen: Score-informed association of sound tracks to players in chamber music performance videos,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2017
2017
Cited alongside, same era.
J. Pu, Y. Panagakis, S. Petridis, and M. Pantic, “Audio-visual object localization and separation using low-rank and sparsity,” 2017
2017
Cited alongside, same era.
A. Torfi, S. M. Iranmanesh, N. M. Nasrabadi, and J. M. Dawson, “Coupled 3d convolutional neural networks for audio-visual recognition,” CoRR , 2017
2017
Cited alongside, same era.
C. Lippert, R. Sabatini, M. C. Maher, E. Y. Kang, S. Lee, O. Arikan, A. Harley, A. Bernal, P. Garst, V. Lavrenko, K. Yocum, T. Wong, M. Zhu, W.-Y. Yang, C. Chang, T. Lu, C. W. H. Lee, B. Hicks, S. Ramakrishnan, H. Tang, C. Xie, J. Piper, S. Brewerton, Y. Turpaz, A. Telenti, R. K. Roby, F. J. Och, and J. C. Venter, “Identification of individuals by trait prediction using whole-genome sequencing data,” Proceedings of the National Academy of Sciences , pp. 10 166–10 171, 2017
2017
Cited alongside, same era.
K. Hoover, S. Chaudhuri, C. Pantofaru, M. Slaney, and I. Sturdy, “Putting a face to the voice: Fusing audio and visual signals across a video to determine speakers,” CoRR , 2017
2017
Cited alongside, same era.
Y. Aytar, C. Vondrick, and A. Torralba, “See, hear, and read: Deep aligned representations,” CoRR , 2017
2017
Cited alongside, same era.
2018
Later among the works it cites.
D. Surís, A. Duarte, A. Salvador, J. Torres, and X. Giró i Nieto, “Cross-modal embeddings for video and audio retrieval,” CoRR , 2018
2018
Later among the works it cites.
A. Nagrani, S. Albanie, and A. Zisserman, “Learnable pins: Cross-modal embeddings for person identity,” CoRR , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
L. Wei, S. Zhang, W. Gao, and Q. Tian, “Person transfer gan to bridge domain gap for person re-identification,” in CVPR , 2018, pp. 79–88
2018
Later among the works it cites.
S.-W. Huang, C.-T. Lin, S.-P. Chen, Y.-Y. Wu, P.-H. Hsu, and S.-H. Lai, “Auggan: Cross domain adaptation with gan-based data augmentation,” in ECCV , 2018, pp. 718–731
2018
Later among the works it cites.
2018
Later among the works it cites.
Y. Qiu and H. Kataoka, “Image generation associated with music data,” in 2018 IEEE Conference on Computer Vision and Pattern Recognition Workshops, CVPR Workshops 2018, Salt Lake City, UT, USA, June 18-22, 2018 , 2018, pp. 2510–2513
2018
Later among the works it cites.
J. Lee, S. Kim, and K. Lee, “Listen to dance: Music-driven choreography generation using autoregressive encoder-decoder network,” CoRR , 2018
2018
Later among the works it cites.
E. Shlizerman, L. Dery, H. Schoen, and I. Kemelmacher-Shlizerman, “Audio to body dynamics,” in Proc. CVPR , 2018
2018
Later among the works it cites.
T. Tang, J. Jia, and H. Mao, “Dance with melody: An lstm-autoencoder approach to music-oriented dance synthesis,” in 2018 ACM Multimedia Conference on Multimedia Conference, MM 2018, Seoul, Republic of Korea, October 22-26, 2018 , 2018, pp. 1598–1606
2018
Later among the works it cites.
N. Yalta, S. Watanabe, K. Nakadai, and T. Ogata, “Weakly supervised deep recurrent neural networks for basic dance step generation,” CoRR , 2018
2018
Later among the works it cites.
S. A. Jalalifar, H. Hasani, and H. Aghajan, “Speech-driven facial reenactment using conditional generative adversarial networks,” CoRR , 2018
2018
Later among the works it cites.
K. Vougioukas, S. Petridis, and M. Pantic, “End-to-end speech-driven facial animation with temporal gans,” in BMVC , 2018
2018
Later among the works it cites.
L. Chen, Z. Li, R. K. Maddox, Z. Duan, and C. Xu, “Lip movements generation at a glance,” CoRR , 2018
2018
Later among the works it cites.
H. Zhou, Y. Liu, Z. Liu, P. Luo, and X. Wang, “Talking face generation by adversarially disentangled audio-visual representation,” CoRR , 2018
2018
Later among the works it cites.
O. Wiles, A. Koepke, and A. Zisserman, “X2face: A network for controlling face generation by using images, audio, and pose codes,” in European Conference on Computer Vision , 2018
2018
Later among the works it cites.
S. Parekh, S. Essid, A. Ozerov, N. Q. Duong, P. Pérez, and G. Richard, “Weakly supervised representation learning for unsynchronized audio-visual events,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops , 2018, pp. 2518–2519
2018
Later among the works it cites.
2018
Later among the works it cites.
A. Owens and A. A. Efros, “Audio-visual scene analysis with self-supervised multisensory features,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 631–648
2018
Later among the works it cites.
N. Alghamdi, S. Maddock, R. Marxer, J. Barker, and G. J. Brown, “A corpus of audio-visual lombard speech with frontal and profile views,” The Journal of the Acoustical Society of America , pp. EL523–EL529, 2018
2018
Later among the works it cites.
S. R. Livingstone and F. A. Russo, “The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,” PloS one , p. e0196391, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
C. Gu, C. Sun, D. A. Ross, C. Vondrick, C. Pantofaru, Y. Li, S. Vijayanarasimhan, G. Toderici, S. Ricco, R. Sukthankar et al. , “Ava: A video dataset of spatio-temporally localized atomic visual actions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 6047–6056
2018
Later among the works it cites.
G. Krishna, C. Tran, J. Yu, and A. H. Tewfik, “Speech recognition with no speech or with noisy speech,” in ICASSP , 2019, pp. 1090–1094
2019
Later among the works it cites.
C. Fu, X. Wu, Y. Hu, H. Huang, and R. He, “Dual variational generation for low-shot heterogeneous face recognition,” NeurIPS , 2019
2019
Later among the works it cites.
T. Karras, S. Laine, and T. Aila, “A style-based generator architecture for generative adversarial networks,” in CVPR , 2019, pp. 4401–4410
2019
Later among the works it cites.
G. Morrone, S. Bergamaschi, L. Pasa, L. Fadiga, V. Tikhanoff, and L. Badino, “Face landmark-based speaker-independent audio-visual speech enhancement in multi-talker environments,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 6900–6904
2019
Later among the works it cites.
H. Zhao, C. Gan, W. Ma, and A. Torralba, “The sound of motions,” CoRR , 2019
2019
Later among the works it cites.
A. Rouditchenko, H. Zhao, C. Gan, J. H. McDermott, and A. Torralba, “Self-supervised audio-visual co-segmentation,” CoRR , 2019
2019
Later among the works it cites.
R. Białobrzeski, M. Kośmider, M. Matuszewski, M. Plata, and A. Rakowski, “Robust bayesian and light neural networks for voice spoofing detection,” Proc. Interspeech 2019 , pp. 1028–1032, 2019
2019
Later among the works it cites.
A. Gomez-Alanis, A. M. Peinado, J. A. Gonzalez, and A. M. Gomez, “A light convolutional gru-rnn deep feature extractor for asv spoofing detection,” Proc. Interspeech 2019 , pp. 1068–1072, 2019
2019
Later among the works it cites.
R. Wang, H. Huang, X. Zhang, J. Ma, and A. Zheng, “A novel distance learning for elastic cross-modal audio-visual matching,” in 2019 IEEE International Conference on Multimedia & Expo Workshops (ICMEW) , 2019, pp. 300–305
2019
Later among the works it cites.
R. Wang, H. Huang, X. Zhang, J. Ma, and A. Zheng, “A novel distance learning for elastic cross-modal audio-visual matching,” in 2019 IEEE International Conference on Multimedia Expo Workshops (ICMEW) , 2019, pp. 300–305
2019
Later among the works it cites.
A. Duarte, F. Roldan, M. Tubau, J. Escur, S. Pascual, A. Salvador, E. Mohedano, K. McGuinness, J. Torres, and X. Giro-i Nieto, “Speech-conditioned face generation using generative adversarial networks,” 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
Z. D. C. X. Lele Chen, Ross K Maddox, “Hierarchical cross-modal talking face generation with dynamic pixel-wise loss,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
B. Li, X. Liu, K. Dinesh, Z. Duan, and G. Sharma, “Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications,” IEEE Transactions on Multimedia , pp. 522–535, 2019
2019
Later among the works it cites.