Fetching the paper…
Reading the bibliography…
Speaker extraction algorithm relies on the speech sample from the target speaker as the reference point to focus its attention.
“Theory and application of digital signal processing,”
AV Oppenheim and R Schafer, · 1978
Earlier work this paper cites.
“The cocktail party phenomenon: A review of research on speech intelligibility in multiple-talker conditions,”
Adelbert W Bronkhorst, · 2000
Earlier work this paper cites.
“Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,”
Antony W Rix, John G Beerends, Michael P Hollier, and Andries P Hekstra, · 2001
Earlier work this paper cites.
“An audio-visual corpus for speech perception and automatic speech recognition,”
Martin Cooke, Jon Barker, Stuart Cunningham, and Xu Shao, · 2006
Earlier work this paper cites.
“Auditory segmentation based on onset and offset analysis,”
Guoning Hu and DeLiang Wang, · 2007
Earlier work this paper cites.
“Monaural sound source separation by nonnegative matrix factorization with temporal continuity and sparseness criteria,”
Tuomas Virtanen, · 2007
Earlier work this paper cites.
“TCD-TIMIT: An audio-visual corpus of continuous speech,”
Naomi Harte and Eoin Gillen, · 2015
Earlier work this paper cites.
“Deep clustering: Discriminative embeddings for segmentation and separation,”
John R Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe, · 2016
Earlier work this paper cites.
“Single channel speech separation with constrained utterance level permutation invariant training using grid lstm,”
Chenglin Xu, Wei Rao, Xiong Xiao, Eng Siong Chng, and Haizhou Li, · 2018
Earlier work this paper cites.
“Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation,”
Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T Freeman, and Michael Rubinstein, · 2018
Cited alongside, same era.
“The Conversation: Deep audio-visual speech enhancement,”
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman, · 2018
Cited alongside, same era.
“Deep audio-visual speech recognition,”
Triantafyllos Afouras, Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman, · 2018
Cited alongside, same era.
“VoxCeleb2: Deep speaker recognition,”
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman, · 2018
Cited alongside, same era.
“LRS3-TED: a large-scale dataset for visual speech recognition,”
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman, · 2018
Cited alongside, same era.
“Multimodal speakerbeam: Single channel target speech extraction with audio-visual speaker clues,”
Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Atsunori Ogawa, and Tomohiro Nakatani, · 2019
Later among the works it cites.
“My lips are concealed: Audio-visual speech enhancement through obstructions,”
T. Afouras, J. S. Chung, and A. Zisserman, · 2019
Later among the works it cites.
“SDR–half-baked or well done?,”
Jonathan Le Roux, Scott Wisdom, Hakan Erdogan, and John R Hershey, · 2019
Later among the works it cites.
“Audio-visual speech separation using i-vectors,”
Luo Yiyu, Wang Jing, Wang Xinyao, Wen Liang, and Wang Lizhong, · 2019
Later among the works it cites.
“SpEx: Multi-scale time domain speaker extraction network,”
Chenglin Xu, Wei Rao, Eng Siong Chng, and Haizhou Li, · 2020
Closest in time.
“Speakerfilter: Deep learning-based target speaker extraction using anchor speech,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Audio-visual scene analysis with self-supervised multisensory features,”
Andrew Owens and Alexei A Efros, · 2018
Cited alongside, same era.
“Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation,”
Yi Luo and Nima Mesgarani, · 2019
Cited alongside, same era.
“VoiceFilter: Targeted voice separation by speaker-conditioned spectrogram masking,”
Quan Wang, Hannah Muckenhirn, Kevin Wilson, Prashant Sridhar, Zelin Wu, John R Hershey, Rif A Saurous, Ron J Weiss, Ye Jia, and Ignacio Lopez Moreno, · 2019
Cited alongside, same era.
“Time domain audio visual speech separation,”
Jian Wu, Yong Xu, Shi-Xiong Zhang, Lian-Wu Chen, Meng Yu, Lei Xie, and Dong Yu, · 2019
Cited alongside, same era.
Shulin He, Hao Li, and Xueliang Zhang, · 2020
Closest in time.
“SpEx+: A complete time domain speaker extraction network,”
Meng Ge, Chenglin Xu, Longbiao Wang, Eng Siong Chng, Jianwu Dang, and Haizhou Li, · 2020
Closest in time.
“AVA active speaker: An audio-visual dataset for active speaker detection,”
Joseph Roth, Sourish Chaudhuri, Ondrej Klejch, Radhika Marvin, Andrew Gallagher, Liat Kaver, Sharadh Ramaswamy, Arkadiusz Stopczynski, Cordelia Schmid, Zhonghua Xi, et al., · 2020
Closest in time.
“FaceFilter: Audio-Visual Speech Separation Using Still Images,”
Soo-Whan Chung, Soyeon Choe, Joon Son Chung, and Hong-Goo Kang, · 2020
Closest in time.