Fetching the paper…
Reading the bibliography…
Streaming end-to-end multi-talker speech recognition aims at transcribing the overlapped speech from conversations or meetings with an all-neural model in a streaming fashion, which is fundamentally different from a modular-based approach that usually cascades the speech separation and the speech recognition models trained independently.
“Robust end-of-utterance detection for real-time speech recognition applications,”
Ramalingam Hariharan, Jula Hakkinen, and Kari Laurila, · 2001
Earlier work this paper cites.
“The kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al., · 2011
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Real-life voice activity detection with lstm recurrent neural networks and an application to hollywood movies,”
Florian Eyben, Felix Weninger, Stefano Squartini, and Björn Schuller, · 2013
Earlier work this paper cites.
“Improvements to the ibm speech activity detection system for the darpa rats program,”
Samuel Thomas, George Saon, Maarten Van Segbroeck, and Shrikanth S Narayanan, · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Neural machine translation of rare words with subword units,” 2016
Rico Sennrich, Barry Haddow, and Alexandra Birch, · 2016
Earlier work this paper cites.
“Permutation invariant training of deep models for speaker-independent multi-talker speech separation,”
Dong Yu, Morten Kolbæk, Zheng-Hua Tan, and Jesper Jensen, · 2017
Earlier work this paper cites.
Takuya Yoshioka, Hakan Erdogan, Zhuo Chen, Xiong Xiao, and Fil Alleva, · 2018
Earlier work this paper cites.
“End-to-end multi-speaker speech recognition,”
Shane Settle, Jonathan Le Roux, Takaaki Hori, Shinji Watanabe, and John R Hershey, · 2018
Earlier work this paper cites.
“End-to-end monaural multi-speaker asr system without pretraining,”
Xuankai Chang, Yanmin Qian, Kai Yu, and Shinji Watanabe, · 2019
Cited alongside, same era.
“Joint endpointing and decoding with end-to-end models,”
Shuo-Yiin Chang, Rohit Prabhavalkar, Yanzhang He, Tara N Sainath, and Gabor Simko, · 2019
Cited alongside, same era.
“Transformer-transducer: End-to-end speech recognition with self-attention,”
Ching-Feng Yeh, Jay Mahadeokar, Kaustubh Kalgaonkar, Yongqiang Wang, Duc Le, Mahaveer Jain, Kjell Schubert, Christian Fuegen, and Michael L Seltzer, · 2019
Cited alongside, same era.
“A comparative study on transformer vs rnn in speech applications,”
Shigeki Karita, Nanxin Chen, Tomoki Hayashi, Takaaki Hori, Hirofumi Inaguma, Ziyan Jiang, Masao Someki, Nelson Enrique Yalta Soplin, Ryuichi Yamamoto, Xiaofei Wang, et al., · 2019
Cited alongside, same era.
“Continuous speech separation: Dataset and analysis,”
Zhuo Chen, Takuya Yoshioka, Liang Lu, Tianyan Zhou, Zhong Meng, Yi Luo, Jian Wu, Xiong Xiao, and Jinyu Li, · 2020
Cited alongside, same era.
“Joint speaker counting, speech recognition, and speaker identification for overlapped speech of any number of speakers,”
Naoyuki Kanda, Yashesh Gaur, Xiaofei Wang, Zhong Meng, Zhuo Chen, Tianyan Zhou, and Takuya Yoshioka, · 2020
Later among the works it cites.
“Exploring transformers for large-scale speech recognition,”
Liang Lu, Changliang Liu, Jinyu Li, and Yifan Gong, · 2020
Later among the works it cites.
“Streaming end-to-end multi-talker speech recognition,”
Liang Lu, Naoyuki Kanda, Jinyu Li, and Yifan Gong, · 2021
Later among the works it cites.
“Streaming multi-speaker asr with rnn-t,”
Ilya Sklyar, Anna Piunova, and Yulan Liu, · 2021
Later among the works it cites.
“Continuous streaming multi-talker asr with dual-path transducers,”
Desh Raj, Liang Lu, Zhuo Chen, Yashesh Gaur, and Jinyu Li, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Serialized output training for end-to-end overlapped speech recognition,”
Naoyuki Kanda, Yashesh Gaur, Xiaofei Wang, Zhong Meng, and Takuya Yoshioka, · 2020
Cited alongside, same era.
“End-to-end multi-talker overlapping speech recognition,”
Anshuman Tripathi, Han Lu, and Hasim Sak, · 2020
Cited alongside, same era.
“Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss,”
Qian Zhang, Han Lu, Hasim Sak, Anshuman Tripathi, Erik McDermott, Stephen Koo, and Shankar Kumar, · 2020
Cited alongside, same era.
“Towards fast and accurate streaming end-to-end ASR,”
Bo Li, Shuo-yiin Chang, Tara N Sainath, Ruoming Pang, Yanzhang He, Trevor Strohman, and Yonghui Wu, · 2020
Cited alongside, same era.
“Developing real-time streaming transformer transducer for speech recognition on large-scale dataset,”
Xie Chen, Yu Wu, Zhenghao Wang, Shujie Liu, and Jinyu Li, · 2021
Later among the works it cites.
“FastEmit: Low-latency streaming ASR with sequence-level emission regularization,”
Jiahui Yu, Chung-Cheng Chiu, Bo Li, Shuo-yiin Chang, Tara N Sainath, Yanzhang He, Arun Narayanan, Wei Han, Anmol Gulati, Yonghui Wu, et al., · 2021
Later among the works it cites.
“Minimum bayes risk training for end-to-end speaker-attributed asr,”
Naoyuki Kanda, Zhong Meng, Liang Lu, Yashesh Gaur, Xiaofei Wang, Zhuo Chen, and Takuya Yoshioka, · 2021
Later among the works it cites.
“Streaming multi-talker speech recognition with joint speaker identification,”
Liang Lu, Naoyuki Kanda, Jinyu Li, and Yifan Gong, · 2021
Later among the works it cites.