Fetching the paper…
Reading the bibliography…
Visual speech, referring to the visual domain of speech, has attracted increasing attention due to its wide applications, such as public security, medical treatment, military defense, and film entertainment.
L. Li, S. Wang, Z. Zhang, Y. Ding, Y. Zheng, X. Yu, and C. Fan, “Write-a-speaker: Text-based emotional and rhythmic talking-head generation,” in
1920
Earlier work this paper cites.
S. E. Eskimez, R. K. Maddox, C. Xu, and Z. Duan, “End-to-end generation of talking faces from noisy speech,” in
1952
Earlier work this paper cites.
H. McGurk and J. MacDonald, “Hearing lips and seeing voices,”
1976
Earlier work this paper cites.
W. Chen, X. Tan, Y. Xia, T. Qin, Y. Wang, and T.-Y. Liu, “Duallip: A system for joint lip reading and generation,” in
1993
Earlier work this paper cites.
C. Bregler, M. Covell, and M. Slaney, “Video rewrite: Driving visual speech with audio,” in
1997
Earlier work this paper cites.
G. Potamianos, H. P. Graf, and E. Cosatto, “An image transform approach for hmm based automatic lipreading,” in
1998
Earlier work this paper cites.
E. S. Ristad and P. N. Yianilos, “Learning string-edit distance,”
1998
Earlier work this paper cites.
K. Kirchho, “Robust speech recognition using articulatory information,” Ph.D. dissertation, Citeseer, 1999
1999
Earlier work this paper cites.
S. Dupont and J. Luettin, “Audio-visual speech modeling for continuous speech recognition,”
2000
Earlier work this paper cites.
T. Chen, “Audiovisual speech processing,”
2001
Earlier work this paper cites.
I. Matthews, T. F. Cootes, J. A. Bangham, S. Cox, and R. Harvey, “Extraction of visual features for lipreading,”
2002
Earlier work this paper cites.
G. Potamianos, C. Neti, G. Gravier, A. Garg, and A. W. Senior, “Recent advances in the automatic recognition of audiovisual speech,”
2003
Earlier work this paper cites.
B. Lee, M. Hasegawa-Johnson, C. Goudeseune, S. Kamdar, S. Borys, M. Liu, and T. Huang, “Avicar: Audio-visual speech corpus in a car environment,” in
2004
Earlier work this paper cites.
S. Fu, R. Gutierrez-Osuna, A. Esposito, P. K. Kakumanu, and O. N. Garcia, “Audio/visual mapping with cross-modal hidden markov models,”
2005
Earlier work this paper cites.
M. Cooke, J. Barker, S. Cunningham, and X. Shao, “An audio-visual corpus for speech perception and automatic speech recognition,”
2006
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in
2006
Earlier work this paper cites.
R. Hadsell, S. Chopra, and Y. LeCun, “Dimensionality reduction by learning an invariant mapping,” in
2006
Earlier work this paper cites.
N. Tye-Murray, M. S. Sommers, and B. Spehar, “Audiovisual integration and lipreading abilities of older adults with normal and impaired hearing,”
2007
Earlier work this paper cites.
L. Xie and Z.-Q. Liu, “Realistic mouth-synching for speech-driven talking face using articulatory modelling,”
2007
Earlier work this paper cites.
N. D. Narvekar and L. J. Karam, “A no-reference perceptual image sharpness metric based on a cumulative probability of blur detection,” in
2009
Earlier work this paper cites.
S. Deena, S. Hou, and A. Galata, “Visual speech synthesis by modelling coarticulation dynamics using a non-parametric switching state-space model,” in
2010
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. Hinton, “ImageNet classification with deep convolutional neural networks,” in
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. Hinton, “ImageNet classification with deep convolutional neural networks,” in
2012
Earlier work this paper cites.
R. Anderson, B. Stenger, V. Wan, and R. Cipolla, “Expressive visual text-to-speech using active appearance models,” in
2013
Earlier work this paper cites.
A. Graves, A.-r. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” in
2013
Earlier work this paper cites.
C. Cao, Y. Weng, S. Zhou, Y. Tong, and K. Zhou, “Facewarehouse: A 3d facial expression database for visual computing,”
2013
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,”
2014
Earlier work this paper cites.
Z. Zhou, G. Zhao, X. Hong, and M. Pietikäinen, “A review of recent advances in visual speech decoding,”
2014
Earlier work this paper cites.
P. Garrido, L. Valgaerts, O. Rehmsen, T. Thormahlen, P. Perez, and C. Theobalt, “Automatic face reenactment,” in
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
M. Mirza and S. Osindero, “Conditional generative adversarial nets,”
2014
Earlier work this paper cites.
Z. Akhtar, C. Micheloni, and G. L. Foresti, “Biometric liveness detection: Challenges and research opportunities,”
2015
Earlier work this paper cites.
P. Garrido, L. Valgaerts, H. Sarmadi, I. Steiner, K. Varanasi, P. Perez, and C. Theobalt, “Vdub: Modifying face video of actors for plausible visual alignment to a dubbed audio track,” in
2015
Earlier work this paper cites.
T. Kim, Y. Yue, S. Taylor, and I. Matthews, “A decision tree framework for spatiotemporal sequence prediction,” in
2015
Earlier work this paper cites.
B. Fan, L. Wang, F. K. Soong, and L. Xie, “Photo-real talking head with deep bidirectional lstm,” in
2015
Earlier work this paper cites.
W. Mattheyses and W. Verhelst, “Audiovisual speech synthesis: An overview of the state-of-the-art,”
2015
Earlier work this paper cites.
I. Anina, Z. Zhou, G. Zhao, and M. Pietikäinen, “Ouluvs2: A multi-view audiovisual database for non-rigid mouth motion analysis,” in
2015
Earlier work this paper cites.
Y. Mroueh, E. Marcheret, and V. Goel, “Deep multimodal learning for audio-visual speech recognition,” in
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large scale image recognition,” in
2015
Earlier work this paper cites.
X. Shi, Z. Chen, H. Wang, D.-Y. Yeung, W.-K. Wong, and W.-c. Woo, “Convolutional lstm network: A machine learning approach for precipitation nowcasting,”
2015
Earlier work this paper cites.
J. S. Chung and A. Zisserman, “Lip reading in the wild,” in
2016
Earlier work this paper cites.
D. He, Y. Xia, T. Qin, L. Wang, N. Yu, T.-Y. Liu, and W.-Y. Ma, “Dual learning for machine translation,”
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
J. S. Chung and A. Zisserman, “Out of time: automated lip sync in the wild,” in
2016
Earlier work this paper cites.
J. Thies, M. Zollhofer, M. Stamminger, C. Theobalt, and M. Nießner, “Face2face: Real-time face capture and reenactment of rgb videos,” in
2016
Earlier work this paper cites.
V. Verkhodanova, A. Ronzhin, I. Kipyatkova, D. Ivanko, A. Karpov, and M. Železnỳ, “Havrus corpus: high-speed recordings of audio-visual russian speech,” in
2016
Earlier work this paper cites.
B. Amos, B. Ludwiczuk, M. Satyanarayanan
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in
2016
Earlier work this paper cites.
A. Gabbay, A. Shamir, and S. Peleg, “Visual speech enhancement,”
2017
Earlier work this paper cites.
A. Ephrat and S. Peleg, “Vid2speech: speech reconstruction from silent video,” in
2017
Earlier work this paper cites.
A. Ephrat, T. Halperin, and S. Peleg, “Improved speech reconstruction from silent video,” in
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
T. Karras, T. Aila, S. Laine, A. Herva, and J. Lehtinen, “Audio-driven facial animation by joint end-to-end learning of pose and emotion,”
2017
Earlier work this paper cites.
J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman, “Lip reading sentences in the wild,” in
2017
Earlier work this paper cites.
A. Nagrani, J. S. Chung, and A. Zisserman, “Voxceleb: a large-scale speaker identification dataset,”
2017
Earlier work this paper cites.
T. Stafylakis and G. Tzimiropoulos, “Combining residual networks with LSTMs for lipreading,” in
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
S. Taylor, T. Kim, Y. Yue, M. Mahler, J. Krahe, A. G. Rodriguez, J. Hodgins, and I. Matthews, “A deep learning approach for generalized speech animation,”
2017
Earlier work this paper cites.
A. Czyzewski, B. Kostek, P. Bratoszewski, J. Kotus, and M. Szykulski, “An audio-visual corpus for multimodal automatic speech recognition,”
2017
Earlier work this paper cites.
S. Suwajanakorn, S. M. Seitz, and I. Kemelmacher-Shlizerman, “Synthesizing obama: learning lip sync from audio,”
2017
Earlier work this paper cites.
A. Fernandez-Lopez, O. Martinez, and F. M. Sukno, “Towards estimating the upper bound of visual-speech recognition: The visual lip-reading feasibility database,” in
2017
Earlier work this paper cites.
A. Koumparoulis, G. Potamianos, Y. Mroueh, and S. J. Rennie, “Exploring roi size in deep learning based lipreading.” in
2017
Earlier work this paper cites.
A. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “Mobilenets: Efficient convolutional neural networks for mobile vision applications,” in
2017
Cited alongside, same era.
G. Huang, Z. Liu, K. Q. Weinberger, and L. van der Maaten, “Densely connected convolutional networks,” in
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Cited alongside, same era.
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in
2017
Cited alongside, same era.
P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros, “Image-to-image translation with conditional adversarial networks,” in
2017
Cited alongside, same era.
H. Liu, Z. Chen, and B. Yang, “Lip graph assisted audio-visual speech recognition using bidirectional synchronous fusion.” in
2020
Later among the works it cites.
B. Martinez, P. Ma, S. Petridis, and M. Pantic, “Lipreading using temporal convolutional networks,” in
2020
Later among the works it cites.
J. Thies, M. Elgharib, A. Tewari, C. Theobalt, and M. Nießner, “Neural voice puppetry: Audio-driven facial reenactment,” in
2020
Later among the works it cites.
K. Wang, Q. Wu, L. Song, Z. Yang, W. Wu, C. Qian, R. He, Y. Qiao, and C. C. Loy, “Mead: A large-scale audio-visual dataset for emotional talking-face generation,” in
2020
Later among the works it cites.
S.-W. Chung, J. S. Chung, and H.-G. Kang, “Perfect match: Self-supervised embeddings for cross-modal retrieval,”
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. X. Pham, S. Cheung, and V. Pavlovic, “Speech-driven 3d facial animation with implicit emotional awareness: a deep learning approach,” in
2017
Cited alongside, same era.
J. S. Chung, A. Jamaludin, and A. Zisserman, “You said that?” in
2017
Cited alongside, same era.
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, P. Nguyen, R. Pang, I. Lopez Moreno, Y. Wu
2018
Cited alongside, same era.
T. Afouras, J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman, “Deep audio-visual speech recognition,”
2018
Cited alongside, same era.
K. Sun, C. Yu, W. Shi, L. Liu, and Y. Shi, “Lip-interact: Improving mobile device interaction with silent speech commands,” in
2018
Cited alongside, same era.
J. S. Chung, A. Nagrani, and A. Zisserman, “Voxceleb2: Deep speaker recognition,”
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2020
Later among the works it cites.
X. Chen, J. Du, and H. Zhang, “Lipreading with densenet and resbi-lstm,”
2020
Later among the works it cites.
J. Yu, S.-X. Zhang, J. Wu, S. Ghorbani, B. Wu, S. Kang, S. Liu, X. Liu, H. Meng, and D. Yu, “Audio-visual recognition of overlapped speech for the lrs2 dataset,” in
2020
Later among the works it cites.
S. Sinha, S. Biswas, and B. Bhowmick, “Identity-preserving realistic talking face generation,” in
2020
Later among the works it cites.
D. Das, S. Biswas, S. Sinha, and B. Bhowmick, “Speech-driven facial animation using cascaded gans for learning of motion and texture,” in
2020
Later among the works it cites.
Y.-A. Chung and J. Glass, “Generative pre-training for speech with autoregressive predictive coding,” in
2020
Later among the works it cites.
2020
Later among the works it cites.
P. Tzirakis, A. Papaioannou, A. Lattas, M. Tarasiou, B. Schuller, and S. Zafeiriou, “Synthesising 3d facial motion from “in-the-wild” speech,” in
2020
Later among the works it cites.
K. Vougioukas, S. Petridis, and M. Pantic, “Realistic speech-driven facial animation with gans,”
2020
Later among the works it cites.
D. Zeng, H. Liu, H. Lin, and S. Ge, “Talking face generation with expression-tailored generative adversarial network,” in
2020
Later among the works it cites.
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” in
2020
Later among the works it cites.
X. Ji, H. Zhou, K. Wang, W. Wu, C. C. Loy, X. Cao, and F. Xu, “Audio-driven emotional video portraits,” in
2021
Later among the works it cites.
A. Haliassos, K. Vougioukas, S. Petridis, and M. Pantic, “Lips don’t lie: A generalisable and robust approach to face forgery detection,” in
2021
Later among the works it cites.
S. Ren, Y. Du, J. Lv, G. Han, and S. He, “Learning from the master: Distilling cross-modal advanced knowledge for lip reading,” in
2021
Later among the works it cites.
2021
Later among the works it cites.
S. Fenghour, D. Chen, K. Guo, B. Li, and P. Xiao, “Deep learning-based automated lip-reading: A survey,”
2021
Later among the works it cites.
C. Sheng, M. Pietikäinen, Q. Tian, and L. Liu, “Cross-modal self-supervised learning for lip reading: When contrastive learning meets adversarial training,” in
2021
Later among the works it cites.
P. Ma, B. Martinez, S. Petridis, and M. Pantic, “Towards practical lipreading with distilled and efficient models,” in
2021
Later among the works it cites.
C. Sheng, X. Zhu, H. Xu, M. Pietikainen, and L. Liu, “Adaptive semantic-spatio-temporal graph convolutional network for lip reading,”
2021
Later among the works it cites.
Z. Zhang, L. Li, Y. Ding, and C. Fan, “Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset,” in
2021
Later among the works it cites.
A. Lahiri, V. Kwatra, C. Frueh, J. Lewis, and C. Bregler, “Lipsync3d: Data-efficient learning of personalized 3d talking faces from video using pose and lighting normalization,” in
2021
Later among the works it cites.
Y. Sun, H. Zhou, Z. Liu, and H. Koike, “Speech2talking-face: Inferring and driving a face with synchronized audio-visual representation,” in
2021
Later among the works it cites.
Y. Guo, K. Chen, S. Liang, Y.-J. Liu, H. Bao, and J. Zhang, “Ad-nerf: Audio driven neural radiance fields for talking head synthesis,” in
2021
Later among the works it cites.
D. Serdyuk, O. Braga, and O. Siohan, “Audio-visual speech recognition is worth
2021
Later among the works it cites.
C. Zhang and H. Zhao, “Lip reading using local-adjacent feature extractor and multi-level feature fusion,” in
2021
Later among the works it cites.
Z. Li, F. Liu, W. Yang, S. Peng, and J. Zhou, “A survey of convolutional neural networks: analysis, applications, and prospects,”
2021
Later among the works it cites.
P. Ma, S. Petridis, and M. Pantic, “End-to-end audio-visual speech recognition with conformers,” in
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. Lu, J. Chai, and X. Cao, “Live speech portraits: real-time photorealistic talking-head animation,”
2021
Later among the works it cites.
X. Yao, O. Fried, K. Fatahalian, and M. Agrawala, “Iterative text-based editing of talking-heads using neural retargeting,”
2021
Later among the works it cites.
H. Wu, J. Jia, H. Wang, Y. Dou, C. Duan, and Q. Deng, “Imitating arbitrary talking style for realistic audio-driven talking face synthesis,” in
2021
Later among the works it cites.
C. Zhang, Y. Zhao, Y. Huang, M. Zeng, S. Ni, M. Budagavi, and X. Guo, “Facial: Synthesizing dynamic talking face with implicit attribute learning,” in
2021
Later among the works it cites.
J. Liu, B. Hui, K. Li, Y. Liu, Y.-K. Lai, Y. Zhang, Y. Liu, and J. Yang, “Geometry-guided dense perspective network for speech-driven facial animation,”
2021
Later among the works it cites.
A. Richard, M. Zollhöfer, Y. Wen, F. De la Torre, and Y. Sheikh, “Meshtalk: 3d face animation from speech using cross-modality disentanglement,” in
2021
Later among the works it cites.
2021
Later among the works it cites.
S. E. Eskimez, Y. Zhang, and Z. Duan, “Speech driven talking face generation from a single image and an emotion condition,”
2021
Later among the works it cites.
H. Zhou, Y. Sun, W. Wu, C. C. Loy, X. Wang, and Z. Liu, “Pose-controllable talking face generation by implicitly modularized audio-visual representation,” in
2021
Later among the works it cites.
S. Chen, Z. Liu, J. Liu, Z. Yan, and L. Wang, “Talking head generation with audio and speech related facial action units,” in
2021
Later among the works it cites.
H. Zhu, H. Huang, Y. Li, A. Zheng, and R. He, “Arbitrary talking face generation via attentional audio-visual coherence learning,” in
2021
Later among the works it cites.
C.-C. Yang, W.-C. Fan, C.-F. Yang, and Y.-C. F. Wang, “Cross-modal mutual learning for audio-visual speech recognition and manipulation,” in
2022
Closest in time.
K. Prajwal, T. Afouras, and A. Zisserman, “Sub-word level lip reading with visual attention,” in
2022
Closest in time.
Z. Ye, M. Xia, R. Yi, J. Zhang, Y.-K. Lai, X. Huang, G. Zhang, and Y.-j. Liu, “Audio-driven talking face video generation with dynamic convolution kernels,”
2022
Closest in time.
V. S. Kadandale, J. F. Montesinos, and G. Haro, “Vocalist: An audio-visual synchronisation model for lips and voices,” 2022
2022
Closest in time.
K. Han, Y. Wang, H. Chen, X. Chen, J. Guo, Z. Liu, Y. Tang, A. Xiao, C. Xu, Y. Xu
2022
Closest in time.
A. Koumparoulis and G. Potamianos, “Accurate and resource-efficient lipreading with efficientnetv2 and transformers,” in
2022
Closest in time.
L. Song, W. Wu, C. Qian, R. He, and C. C. Loy, “Everybody’s talkin’: Let me talk as you want,”
2022
Closest in time.
Y. Fan, Z. Lin, J. Saito, W. Wang, and T. Komura, “Faceformer: Speech-driven 3d facial animation with transformers,” in
2022
Closest in time.
B. Liang, Y. Pan, Z. Guo, H. Zhou, Z. Hong, X. Han, J. Han, J. Liu, E. Ding, and J. Wang, “Expressive talking head generation with granular audio-visual control,” in
2022
Closest in time.
S. Wang, L. Li, Y. Ding, and X. Yu, “One-shot talking face generation from single-speaker audio-visual correlation learning,” in
2022
Closest in time.
S. Shen, W. Li, Z. Zhu, Y. Duan, J. Zhou, and J. Lu, “Learning dynamic facial radiance fields for few-shot talking head synthesis,” in
2022
Closest in time.
2022
Closest in time.
X. Liu, Y. Xu, Q. Wu, H. Zhou, W. Wu, and B. Zhou, “Semantic-aware implicit neural audio-driven video portrait generation,” in
2022
Closest in time.
2023
Closest in time.
Y. Liu, L. Lin, Y. Fei, Z. Changyin, and L. Yu, “Moda: Mapping-once audio-driven portrait animation with dual attentions,” in
2023
Closest in time.
W. Zhang, X. Cun, X. Wang, Y. Zhang, X. Shen, Y. Guo, Y. Shan, and F. Wang, “Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation,” in
2023
Closest in time.
J. Xing, M. Xia, Y. Zhang, X. Cun, J. Wang, and T.-T. Wong, “Codetalker: Speech-driven 3d facial animation with discrete motion prior,” in
2023
Closest in time.
Z. Ye, Z. Jiang, Y. Ren, J. Liu, J. He, and Z. Zhao, “Geneface: Generalized and high-fidelity audio-driven 3d talking face synthesis,” in
2023
Closest in time.