Fetching the paper…
Reading the bibliography…
This paper presents an audio-visual approach for voice separation which produces state-of-the-art results at a low latency in two scenarios: speech and singing voice.
Cherry, E.C.: Some experiments on the recognition of speech, with one and with two ears. The Journal of the acoustical society of America 25
1953
Earlier work this paper cites.
Kabsch, W.: A discussion of the solution for the best rotation to relate two sets of vectors. Acta Crystallographica Section A 34
1978
Earlier work this paper cites.
Vincent, E., Gribonval, R., Fevotte, C.: Performance measurement in blind audio source separation. IEEE Transactions on Audio, Speech, and Language Processing 14
2005
Earlier work this paper cites.
Ma, W.J., Zhou, X., Ross, L.A., Foxe, J.J., Parra, L.C.: Lip-reading aids word recognition most in moderate noise: a bayesian explanation using high-dimensional feature space. PloS one 4
2009
Earlier work this paper cites.
Golumbic, E.Z., Cogan, G.B., Schroeder, C.E., Poeppel, D.: Visual input enhances selective speech envelope tracking in auditory cortex at a “cocktail party”. The Journal of Neuroscience 33
2013
Earlier work this paper cites.
Williamson, D.S., Wang, Y., Wang, D.: Complex ratio masking for monaural speech separation. IEEE/ACM transactions on audio, speech, and language processing 24
2015
Earlier work this paper cites.
Grais, E.M., Roma, G., Simpson, A.J., Plumbley, M.: Combining mask estimates for single channel audio source separation using deep neural networks. In: Interspeech (2016)
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Gemmeke, J.F., Ellis, D.P., Freedman, D., Jansen, A., Lawrence, W., Moore, R.C., Plakal, M., Ritter, M.: Audio set: An ontology and human-labeled dataset for audio events. In: IEEE Int. Conf. on Acoustics, Speech and Signal Processing (2017)
2017
Earlier work this paper cites.
Rafii, Z., Liutkus, A., Stöter, F.R., Mimilakis, S.I., Bittner, R.: The MUSDB18 corpus for music separation (Dec 2017). https://doi.org/10.5281/zenodo.1117372, https://doi.org/10.5281/zenodo.1117372
2017
Earlier work this paper cites.
Afouras, T., Chung, J.S., Zisserman, A.: The conversation: Deep audio-visual speech enhancement. In: Interspeech (2018)
2018
Earlier work this paper cites.
Chung, J.S., Nagrani, A., Zisserman, A.: Voxceleb2: Deep speaker recognition. In: Interspeech (2018)
2018
Earlier work this paper cites.
Ephrat, A., Mosseri, I., Lang, O., Dekel, T., Wilson, K., Hassidim, A., Freeman, W.T., Rubinstein, M.: Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation. In: SIGGRAPH (2018)
2018
Earlier work this paper cites.
Gabbay, A., Shamir, A., Peleg, S.: Visual speech enhancement. In: Interspeech. pp. 1170–1174. ISCA (2018)
2018
Cited alongside, same era.
Hou, J.C., Wang, S.S., Lai, Y.H., Tsao, Y., Chang, H.W., Wang, H.M.: Audio-visual speech enhancement using multimodal deep convolutional neural networks. IEEE Transactions on Emerging Topics in Computational Intelligence 2
2018
Cited alongside, same era.
Yan, S., Xiong, Y., Lin, D.: Spatial temporal graph convolutional networks for skeleton-based action recognition. In: Proceedings of the AAAI conference on artificial intelligence. vol. 32 (2018)
2018
Cited alongside, same era.
Zhao, H., Gan, C., Rouditchenko, A., Vondrick, C., McDermott, J., Torralba, A.: The sound of pixels. In: Proceedings of the European conference on computer vision (ECCV). pp. 570–586 (2018)
2018
Cited alongside, same era.
Li, C., Qian, Y.: Deep audio-visual speech separation with attention mechanism. In: ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 7314–7318 (2020). https://doi.org/10.1109/ICASSP40776.2020.9054180
2020
Later among the works it cites.
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., Sutskever, I.: Deep double descent: where bigger models and more data hurt. Journal of Statistical Mechanics: Theory and Experiment 2021
2020
Later among the works it cites.
Sun, Z., Wang, Y., Cao, L.: An attention based speaker-independent audio-visual deep learning model for speech enhancement. In: International Conference on Multimedia Modeling. pp. 722–728. Springer (2020)
2020
Later among the works it cites.
Chen, H., Xie, W., Afouras, T., Nagrani, A., Vedaldi, A., Zisserman, A.: Audio-visual synchronisation in the wild. In: 32nd British Machine Vision Conference, BMVC (2021)
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gao, R., Grauman, K.: Co-separating sounds of visual objects. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 3879–3888 (2019)
2019
Cited alongside, same era.
Michelsanti, D., Tan, Z.H., Sigurdsson, S., Jensen, J.: On training targets and objective functions for deep-learning-based audio-visual speech enhancement. In: ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 8077–8081. IEEE (2019)
2019
Cited alongside, same era.
Morrone, G., Bergamaschi, S., Pasa, L., Fadiga, L., Tikhanoff, V., Badino, L.: Face landmark-based speaker-independent audio-visual speech enhancement in multi-talker environments. In: ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 6900–6904. IEEE (2019)
2019
Cited alongside, same era.
Wu, J., Xu, Y., Zhang, S.X., Chen, L.W., Yu, M., Xie, L., Yu, D.: Time domain audio visual speech separation. In: 2019 IEEE automatic speech recognition and understanding workshop (ASRU). pp. 667–673. IEEE (2019)
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Zhao, H., Gan, C., Ma, W.C., Torralba, A.: The sound of motions. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 1735–1744 (2019)
2019
Cited alongside, same era.
Chuang, S.Y., Tsao, Y., Lo, C.C., Wang, H.M.: Lite audio-visual speech enhancement. In: Proc. Interspeech 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Gao, R., Grauman, K.: Visualvoice: Audio-visual speech separation with cross-modal consistency. In: CVPR (2021)
2021
Later among the works it cites.
Makishima, N., Ihori, M., Takashima, A., Tanaka, T., Orihashi, S., Masumura, R.: Audio-visual speech separation using cross-modal correspondence loss. In: ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 6673–6677. IEEE (2021)
2021
Later among the works it cites.
Michelsanti, D., Tan, Z.H., Zhang, S.X., Xu, Y., Yu, M., Yu, D., Jensen, J.: An overview of deep-learning-based audio-visual speech enhancement and separation. IEEE/ACM Transactions on Audio, Speech, and Language Processing 29
2021
Later among the works it cites.
Montesinos, J.F., Kadandale, V.S., Haro, G.: A cappella: Audio-visual singing voice separation. In: 32nd British Machine Vision Conference, BMVC (2021)
2021
Later among the works it cites.
Sadeghi, M., Alameda-Pineda, X.: Mixture of inference networks for vae-based audio-visual speech enhancement. IEEE Transactions on Signal Processing 69
2021
Later among the works it cites.
Sato, H., Ochiai, T., Kinoshita, K., Delcroix, M., Nakatani, T., Araki, S.: Multimodal attention fusion for target speaker extraction. In: 2021 IEEE Spoken Language Technology Workshop (SLT). pp. 778–784. IEEE (2021)
2021
Later among the works it cites.
Slizovskaia, O., Haro, G., Gómez, E.: Conditioned source separation for musical instrument performances. IEEE/ACM Transactions on Audio, Speech, and Language Processing 29
2021
Later among the works it cites.
Truong, T.D., Duong, C.N., Pham, H.A., Raj, B., Le, N., Luu, K., et al.: The right to talk: An audio-visual transformer approach. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 1105–1114 (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.