Fetching the paper…
Reading the bibliography…
This paper investigates the possibility of extracting a target sentence from multi-talker speech using only a keyword as input.
“Deep neural networks for single-channel multi-talker speech recognition,”
Chao Weng, Dong Yu, Michael L. Seltzer, and Jasha Droppo, · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Achieving human parity in conversational speech recognition,”
Wayne Xiong, Jasha Droppo, Xuedong Huang, Frank Seide, Mike Seltzer, Andreas Stolcke, Dong Yu, and Geoffrey Zweig, · 2016
Earlier work this paper cites.
“Progressive joint modeling in unsupervised single-channel overlapped speech recognition,”
Zhehuai Chen, Jasha Droppo, Jinyu Li, and Wayne Xiong, · 2017
Earlier work this paper cites.
“Recognizing multi-talker speech with permutation invariant training,”
Dong Yu, Xuankai Chang, and Yanmin Qian, · 2017
Earlier work this paper cites.
“Permutation invariant training of deep models for speaker-independent multi-talker speech separation,”
Dong Yu, Morten Kolbæk, Zheng-Hua Tan, and Jesper Jensen, · 2017
Earlier work this paper cites.
“Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,”
Morten Kolbæk, Dong Yu, Zheng-Hua Tan, and Jesper Jensen, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“End-to-end multi-speaker speech recognition,”
Shane Settle, Jonathan Le Roux, Takaaki Hori, Shinji Watanabe, and John R Hershey, · 2018
Cited alongside, same era.
“Single-channel multi-talker speech recognition with permutation invariant training,”
Yanmin Qian, Xuankai Chang, and Dong Yu, · 2018
Cited alongside, same era.
“Voicefilter: Targeted voice separation by speaker-conditioned spectrogram masking,”
Quan Wang, Hannah Muckenhirn, Kevin Wilson, Prashant Sridhar, Zelin Wu, John R Hershey, Rif A Saurous, Ron J Weiss, Ye Jia, and Ignacio Lopez Moreno, · 2019
Cited alongside, same era.
“Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,”
Kateřina Žmolíková, Marc Delcroix, Keisuke Kinoshita, Tsubasa Ochiai, Tomohiro Nakatani, Lukáš Burget, and Jan Černockỳ, · 2019
Cited alongside, same era.
“Neural mechanisms underlying concurrent listening of simultaneous speech,”
Natasha Yuriko Santos Kawata, Teruo Hashimoto, and Ryuta Kawashima, · 2020
Cited alongside, same era.
“Investigation of end-to-end speaker-attributed asr for continuous multi-talker recordings,”
Naoyuki Kanda, Xuankai Chang, Yashesh Gaur, Xiaofei Wang, Zhong Meng, Zhuo Chen, and Takuya Yoshioka, · 2021
Later among the works it cites.
“Hypothesis stitcher for end-to-end speaker-attributed asr on long-form multi-talker recordings,”
Xuankai Chang, Naoyuki Kanda, Yashesh Gaur, Xiaofei Wang, Zhong Meng, and Takuya Yoshioka, · 2021
Later among the works it cites.
“Minimum bayes risk training for end-to-end speaker-attributed asr,”
Naoyuki Kanda, Zhong Meng, Liang Lu, Yashesh Gaur, Xiaofei Wang, Zhuo Chen, and Takuya Yoshioka, · 2021
Later among the works it cites.
“A review of speaker diarization: Recent advances with deep learning,”
Tae Jin Park, Naoyuki Kanda, Dimitrios Dimitriadis, Kyu J Han, Shinji Watanabe, and Shrikanth Narayanan, · 2022
Later among the works it cites.
“Streaming target-speaker asr with neural transducer,”
Takafumi Moriya, Hiroshi Sato, Tsubasa Ochiai, Marc Delcroix, and Takahiro Shinozaki, · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Serialized output training for end-to-end overlapped speech recognition,”
Naoyuki Kanda, Yashesh Gaur, Xiaofei Wang, Zhong Meng, and Takuya Yoshioka, · 2020
Cited alongside, same era.
“Investigation of practical aspects of single channel speech separation for asr,”
Jian Wu, Zhuo Chen, Sanyuan Chen, Yu Wu, Takuya Yoshioka, Naoyuki Kanda, Shujie Liu, and Jinyu Li, · 2021
Cited alongside, same era.
“Streaming end-to-end multi-talker speech recognition,”
Liang Lu, Naoyuki Kanda, Jinyu Li, and Yifan Gong, · 2021
Cited alongside, same era.
“Conformer-based target-speaker automatic speech recognition for single-channel audio,”
Yang Zhang, Krishna C Puvvada, Vitaly Lavrukhin, and Boris Ginsburg, · 2023
Closest in time.
“A sidecar separator can convert a single-talker speech recognition system to a multi-talker one,”
Lingwei Meng, Jiawen Kang, Mingyu Cui, Yuejiao Wang, Xixin Wu, and Helen Meng, · 2023
Closest in time.