Fetching the paper…
Reading the bibliography…
In this paper, we address the problem of lip-voice synchronisation in videos containing human face and voice.
J. Hershey and J. Movellan, “Audio vision: Using audio-visual synchrony to locate sounds,” Advances in neural information processing systems , vol. 12, 1999
1999
Earlier work this paper cites.
E. Vincent, R. Gribonval, and C. Févotte, “Performance measurement in blind audio source separation,” IEEE trans. on audio, speech, and language process. , vol. 14, no. 4, pp. 1462–1469, 2006
2006
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in International Conference on Medical image computing and computer-assisted intervention . Springer, 2015, pp. 234–241
2015
Earlier work this paper cites.
J. S. Chung and A. Zisserman, “Out of time: automated lip sync in the wild,” in Asian conference on computer vision . Springer, 2016, pp. 251–263
2016
Earlier work this paper cites.
A. Owens and A. A. Efros, “Audio-visual scene analysis with self-supervised multisensory features,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 631–648
2018
Earlier work this paper cites.
T. Afouras, J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman, “Deep audio-visual speech recognition,” IEEE transactions on pattern analysis and machine intelligence , 2018
2018
Earlier work this paper cites.
B. Korbar, D. Tran, and L. Torresani, “Cooperative learning of audio and video models from self-supervised synchronization,” Advances in Neural Information Processing Systems , vol. 31, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri, “A closer look at spatiotemporal convolutions for action recognition,” in Proceedings of the IEEE conference on Computer Vision and Pattern Recognition , 2018, pp. 6450–6459
2018
Cited alongside, same era.
S.-W. Chung, J. S. Chung, and H.-G. Kang, “Perfect match: Improved cross-modal embeddings for audio-visual synchronisation,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 3965–3969
2019
Cited alongside, same era.
Y.-H. H. Tsai, S. Bai, P. P. Liang, J. Z. Kolter, L.-P. Morency, and R. Salakhutdinov, “Multimodal transformer for unaligned multimodal language sequences,” in Proceedings of the conference. Association for Computational Linguistics. Meeting , vol. 2019. NIH Public Access, 2019, p. 6558
2019
Cited alongside, same era.
S. W. Chung, J. S. Chung, and H.-G. Kang, “Perfect match: Self-supervised embeddings for cross-modal retrieval,” IEEE Journal of Selected Topics in Signal Processing , vol. 14, no. 3, pp. 568–576, 2020
J. F. Montesinos, V. S. Kadandale, and G. Haro, “A cappella: Audio-visual singing voice separation,” in 32nd British Machine Vision Conference, BMVC , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. J. Kim, H. S. Heo, S.-W. Chung, and B.-J. Lee, “End-to-end lip synchronisation based on pattern classification,” in 2021 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2021, pp. 598–605
2021
Later among the works it cites.
T.-D. Truong, C. N. Duong, H. A. Pham, B. Raj, N. Le, K. Luu et al. , “The right to talk: An audio-visual transformer approach,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 1105–1114
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
Y. Shalev and L. Wolf, “End to end lip synchronization with a temporal autoencoder,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2020, pp. 341–350
2020
Cited alongside, same era.
K. Prajwal, R. Mukhopadhyay, V. P. Namboodiri, and C. Jawahar, “A lip sync expert is all you need for speech to lip generation in the wild,” in Proceedings of the 28th ACM International Conference on Multimedia , 2020, pp. 484–492
2020
Cited alongside, same era.
T. Afouras, A. Owens, J. S. Chung, and A. Zisserman, “Self-supervised learning of audio-visual objects from video,” in European Conference on Computer Vision . Springer, 2020, pp. 208–224
2020
Cited alongside, same era.
Y.-B. Lin and Y.-C. F. Wang, “Audiovisual transformer with instance attention for audio-visual event localization,” in Proceedings of the Asian Conference on Computer Vision , 2020
2020
Cited alongside, same era.
P. Ma, S. Petridis, and M. Pantic, “End-to-end audio-visual speech recognition with conformers,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 7613–7617
2021
Later among the works it cites.
L. Zhu and E. Rahtu, “Visually guided sound source separation and localization using self-supervised motion representations,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2022, pp. 1289–1299
2022
Closest in time.
Z. Pan, R. Tao, C. Xu, and H. Li, “Selective listening by synchronizing speech with lips,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2022
2022
Closest in time.