Fetching the paper…
Reading the bibliography…
VoiceFilter-Lite is a speaker-conditioned voice separation model that plays a crucial role in improving speech recognition and speaker verification by suppressing overlapping speech from non-target speakers.
“Multi-style training for robust isolated-word speech recognition,”
Richard Lippmann, Edward Martin, and D Paul, · 1987
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“On the efficient representation and execution of deep acoustic models,”
Raziel Alvarez, Rohit Prabhavalkar, and Anton Bakhtin, · 2016
Earlier work this paper cites.
“Speaker-aware neural network based beamformer for speaker extraction in speech mixtures,”
Kateřina Žmolíková, Marc Delcroix, Keisuke Kinoshita, Takuya Higuchi, Atsunori Ogawa, and Tomohiro Nakatani, · 2017
Earlier work this paper cites.
“Learning speaker representation for neural network based multichannel speaker extraction,”
Kateřina Žmolíková, Marc Delcroix, Keisuke Kinoshita, Takuya Higuchi, Atsunori Ogawa, and Tomohiro Nakatani, · 2017
Earlier work this paper cites.
“Tomato, tomahto. Google Home now supports multiple users,” Google Assistant Blog, 2017
Yury Pinsky, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Dynamic layer normalization for adaptive neural acoustic modeling in speech recognition,”
Taesup Kim, Inchul Song, and Yoshua Bengio, · 2017
Earlier work this paper cites.
“A study on data augmentation of reverberant speech for robust speech recognition,”
Tom Ko, Vijayaditya Peddinti, Daniel Povey, Michael L Seltzer, and Sanjeev Khudanpur, · 2017
Earlier work this paper cites.
“Generation of large-scale simulated utterances in virtual rooms to train deep-neural networks for far-field speech recognition in Google Home,”
Chanwoo Kim, Ananya Misra, Kean Chin, Thad Hughes, Arun Narayanan, Tara Sainath, and Michiel Bacchiani, · 2017
Earlier work this paper cites.
“Deep extractor network for target speaker recovery from single channel speech mixtures,”
Jun Wang, Jie Chen, Dan Su, Lianwu Chen, Meng Yu, Yanmin Qian, and Dong Yu, · 2018
Cited alongside, same era.
“Single channel target speaker extraction and recognition with speaker beam,”
Marc Delcroix, Kateřina Žmolíková, Keisuke Kinoshita, Atsunori Ogawa, and Tomohiro Nakatani, · 2018
Cited alongside, same era.
“FiLM: Visual reasoning with a general conditioning layer,”
Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville, · 2018
Cited alongside, same era.
“Generalized end-to-end loss for speaker verification,”
Li Wan, Quan Wang, Alan Papir, and Ignacio Lopez Moreno, · 2018
Cited alongside, same era.
“VoiceFilter: Targeted voice separation by speaker-conditioned spectrogram masking,”
Quan Wang, Hannah Muckenhirn, Kevin Wilson, Prashant Sridhar, Zelin Wu, John R. Hershey, Rif A. Saurous, Ron J. Weiss, Ye Jia, and Ignacio Lopez Moreno, · 2019
Cited alongside, same era.
“More ways to fine tune Google Assistant for you,” Google Assistant Blog, April 2020
Natasha Jensen, · 2020
Later among the works it cites.
“Version control of speaker recognition systems,”
Quan Wang and Ignacio Lopez Moreno, · 2020
Later among the works it cites.
“Optimizing speech recognition for the edge,”
Yuan Shangguan, Jian Li, Qiao Liang, Raziel Alvarez, and Ian McGraw, · 2020
Later among the works it cites.
“Adaptation algorithms for neural network-based speech recognition: An overview,”
Peter Bell, Joachim Fainberg, Ondrej Klejch, Jinyu Li, Steve Renals, and Pawel Swietojanski, · 2021
Later among the works it cites.
“Improving rnn transducer with target speaker extraction and neural uncertainty estimation,”
Jiatong Shi, Chunlei Zhang, Chao Weng, Shinji Watanabe, Meng Yu, and Dong Yu, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Streaming end-to-end speech recognition for mobile devices,”
Yanzhang He, Tara N Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Yonghui Wu, Ruoming Pang, et al., · 2019
Cited alongside, same era.
“End-to-end multi-speaker speech recognition using speaker embeddings and transfer learning,”
Pavel Denisov and Ngoc Thang Vu, · 2019
Cited alongside, same era.
“VoiceFilter-Lite: Streaming targeted voice separation for on-device speech recognition,”
Quan Wang, Ignacio Lopez Moreno, Mert Saglam, Kevin Wilson, Alan Chiao, Renjie Liu, Yanzhang He, Wei Li, Jason Pelecanos, Marily Nika, and Alexander Gruenstein, · 2020
Cited alongside, same era.
“Spex: Multi-scale time domain speaker extraction network,”
Chenglin Xu, Wei Rao, Eng Siong Chng, and Haizhou Li, · 2020
Cited alongside, same era.
“Personal VAD: Speaker-conditioned voice activity detection,”
Shaojin Ding, Quan Wang, Shuo-yiin Chang, Li Wan, and Ignacio Lopez Moreno, · 2020
Cited alongside, same era.
“Joint speaker counting, speech recognition, and speaker identification for overlapped speech of any number of speakers,”
Naoyuki Kanda, Yashesh Gaur, Xiaofei Wang, Zhong Meng, Zhuo Chen, Tianyan Zhou, and Takuya Yoshioka, · 2020
Cited alongside, same era.
“Continuous speech separation using speaker inventory for long multi-talker recording,”
Cong Han, Yi Luo, Chenda Li, Tianyan Zhou, Keisuke Kinoshita, Shinji Watanabe, Marc Delcroix, Hakan Erdogan, John R. Hershey, Nima Mesgarani, and Zhuo Chen, · 2020
Cited alongside, same era.
“Multi-user VoiceFilter-Lite via attentive speaker embedding,”
Rajeev Rikhye, Quan Wang, Qiao Liang, Yanzhang He, and Ian McGraw, · 2021
Later among the works it cites.
“Personalized keyphrase detection using speaker and environment information,”
Rajeev Rikhye, Quan Wang, Qiao Liang, Yanzhang He, Ding Zhao, Arun Narayanan, Ian McGraw, et al., · 2021
Later among the works it cites.
“A conformer-based asr frontend for joint acoustic echo cancellation, speech enhancement and speech separation,”
Tom O’Malley, Arun Narayanan, Quan Wang, Alex Park, James Walker, and Nathan Howard, · 2021
Later among the works it cites.
“Cross-attention conformer for context modeling in speech enhancement for asr,”
Arun Narayanan, Chung-Cheng Chiu, Tom O’Malley, Quan Wang, and Yanzhang He, · 2021
Later among the works it cites.
“Dr-Vectors: Decision residual networks and an improved loss for speaker recognition,”
Jason Pelecanos, Quan Wang, and Ignacio Lopez Moreno, · 2021
Later among the works it cites.
Midia Yousefi and John HL Hanse, · 2021
Later among the works it cites.