Fetching the paper…
Reading the bibliography…
While a streaming voice assistant system has been used in many applications, this system typically focuses on unnatural, one-shot interactions assuming input from a single voice query without hesitation or disfluency.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Japanese and Korean voice search,”
M. Schuster and K. Nakajima, · 2012
Earlier work this paper cites.
“Improvements to the ibm speech activity detection system for the darpa rats program,”
S. Thomas, G. Saon, M. V. Segbroeck, and S. Naranyanan, · 2015
Earlier work this paper cites.
“Lower frame rate neural network acoustic models,”
G. Pundak and T. N. Sainath, · 2016
Earlier work this paper cites.
“Endpoint detection using grid long short-term memory networks for streaming speech recognition.,”
Shuo-Yiin Chang, Bo Li, Tara N Sainath, Gabor Simko, and Carolina Parada, · 2017
Earlier work this paper cites.
“Turn-taking estimation model based on joint embedding of lexical and prosodic contents,”
C. Liu, C. Ishi, and H. Ishiguro, · 2017
Earlier work this paper cites.
“Towards deep end-of-turn prediction for situated spoken dialogue systems,”
Julian Hough Angelika Maier and and David Schlangen, · 2017
Earlier work this paper cites.
“Generated of large-scale simulated utterances in virtual rooms to train deep-neural networks for far-field speech recognition in google home,”
C. Kim, A. Misra, K. Chin, T. Hughes, A. Narayanan, T. N. Sainath, and M. Bacchiani, · 2017
Earlier work this paper cites.
“Combining acoustic embeddings and decoding features for end-of-utterance detection in real-time far-field speech recognition systems,”
R. Maas, A. Rastrow, C. Ma, G. Lan, K. Goehner, G. Tiwari, S. Joseph, and B. Hoffmeister, · 2018
Earlier work this paper cites.
“Evaluation of real-time deep learning turn-taking models for multiple dialogue scenarios,”
D. Lala, K. Inoue, and T. Kawahara, · 2018
Cited alongside, same era.
“Prediction of turn-taking using multitask learning with prediction of backchannels and fillers,”
R. Masumura, T. Asami, H. Masataki, R. Ishii, and R. Higashinaka, · 2018
Cited alongside, same era.
“State-of-the-art Speech Recognition With Sequence-to-Sequence Models,”
C.-C. Chiu, T. N. Sainath, Y. Wu, et al., · 2018
Cited alongside, same era.
“Joint endpointing and decoding with end-to-end models,”
Shuo-Yiin Chang, Rohit Prabhavalkar, Yanzhang He, Tara N Sainath, and Gabor Simko, · 2019
Cited alongside, same era.
“A unified endpointer using multitask and multidomain training,”
Shuo-Yiin Chang, Bo Li, and Gabor Simko, · 2019
Cited alongside, same era.
“Analysis of effect and timing of fillers in natural turn-taking,”
“Towards fast and accurate streaming end-to-end asr,”
Bo Li, Shuo-yiin Chang, Tara N Sainath, Ruoming Pang, Yanzhang He, Trevor Strohman, and Yonghui Wu, · 2020
Later among the works it cites.
“On the Comparison of Popular End-to-End Models for Large Scale Speech Recognition,”
J. Li, Y. Wu, Y. Gaur, et al., · 2020
Later among the works it cites.
“A new training pipeline for an improved neural transducer,”
A. Zeyer, A. Merboldt, R. Schlüter, and H. Ney, · 2020
Later among the works it cites.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al., · 2020
Later among the works it cites.
“Disfluency detection with unlabeled data and small bert models,”
J. C. Rocholl, V. Zayats, D. D. Walker, N. B. Murad, A. Schneider, and D. J. Liebling, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Lala, S. Nakamura, and T. Kawahara, · 2019
Cited alongside, same era.
“Streaming End-to-end Speech Recognition For Mobile Devices,”
Y. He, T. N. Sainath, R. Prabhavalkar, et al., · 2019
Cited alongside, same era.
“Improving RNN transducer modeling for end-to-end speech recognition,”
J. Li, R. Zhao, H. Hu, and Y. Gong, · 2019
Cited alongside, same era.
“Tied & reduced rnn-t decoder,”
R. Botros and T.N. Sainath, · 2021
Later among the works it cites.
“Fastemit: Low-latency streaming asr with sequence-level emission regularization,”
Jiahui Yu, Chung-Cheng Chiu, Bo Li, Shuo-yiin Chang, Tara N Sainath, Yanzhang He, Arun Narayanan, Wei Han, Anmol Gulati, Yonghui Wu, et al., · 2021
Later among the works it cites.
“An efficient streaming non-recurrent on-device end-to-end model with improvements to rare-word modeling,”
Arun Narayanan Tara N. Sainath, Yanzhang He et al., · 2021
Later among the works it cites.