Fetching the paper…
Reading the bibliography…
We present JOIST, an algorithm to train a streaming, cascaded, encoder end-to-end (E2E) model with both speech-text paired inputs, and text-only unpaired inputs.
“A Recursive Algorithm for The Forced Alignment of Very Long Audio Segments,”
P. J. Moreno, C. F. Joerg, J.-M. Van Thong, and O. Glickman, · 1998
Earlier work this paper cites.
“Phone Recognition using Restricted Boltzmann Machines,”
A. Mohamed and G. Hinton, · 2010
Earlier work this paper cites.
“Making Deep Belief Networks effective for large vocabulary continuous speech recognition,”
T. N. Sainath, B. Kingsbury, B. Ramabhadran, P. Fousek, P. Novak, and A. Mohamed, · 2011
Earlier work this paper cites.
“Conversational speech transcription using context-dependent deep neural networks,”
F. Seide, G. Li, and D. Yu, · 2011
Earlier work this paper cites.
“Bayesian Language Model Interpolation for Mobile Speech Input,”
C. Allauzen and M. Riley, · 2011
Earlier work this paper cites.
“Speech Recognition with Deep Neural Networks,”
A. Graves, A.-R. Mohamed, and G. Hinton, · 2012
Earlier work this paper cites.
“Japanese and Korean voice search,”
M. Schuster and K. Nakajima, · 2012
Earlier work this paper cites.
“Improving Wideband Speech Rcognition using Mixed-bandwidth Training Data in CD-DNN-HMM,”
J. Li, D. Yu, J. Huang, and Y. Gong, · 2012
Earlier work this paper cites.
“Large Scale Deep Neural Network Acoustic Modeling with Semi-supervised Training Data for YouTube Video Transcription,”
H. Liao, E. McDermott, and A. Senior, · 2013
Earlier work this paper cites.
“Feature learning in deep neural networks-studies on speech recognition tasks,”
D. Yu, M. L. Seltzer, J. Li, et al., · 2013
Earlier work this paper cites.
“Long short-term memory recurrent neural network architectures for large scale acoustic modeling,”
H. Sak, A. Senior, and F. Beaufays, · 2014
Earlier work this paper cites.
“Backoff Inspired Features for Maximum Entropy Language Models,”
F. Biadsy, K. Hall, P. J. Moreno, and B. Roark, · 2014
Earlier work this paper cites.
“Librispeech: An ASR Corpus based on Public Domain Audio Books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“Lower frame rate neural network acoustic models,”
G. Pundak and T. N. Sainath, · 2016
Earlier work this paper cites.
“Recent Advances in Google Real-time HMM-driven Unit Selection Synthesizer,”
X. Gonzalvo, S. Tazari, C. Chan, et al., · 2016
Earlier work this paper cites.
“Joint CTC-attention based end-to-end speech recognition using multi-task learning,”
S. Kim, T. Hori, and S. Watanabe, · 2017
Earlier work this paper cites.
“Towards Better Decoding and Language Model Integration in Sequence to Sequence Models,”
J. Chorowski and N. Jaitly, · 2017
Earlier work this paper cites.
“Listening While Speaking: Speech Chain by Deep Learning,”
A. Tjandra, S. Sakti, and S. Nakamura, · 2017
Earlier work this paper cites.
“A Comparison of Sequence-to-sequence Models for Speech Recognition,”
R. Prabhavalkar, K. Rao, T. N. Sainath, et al., · 2017
Earlier work this paper cites.
“Generation of Large-Scale Simulated Utterances in Virtual Rooms to Train Deep-Neural Networks for Far-Field Speech Recognition in Google Home,”
C. Kim, A. Misra, K. Chin, et al., · 2017
Earlier work this paper cites.
“State-of-the-art Speech Recognition With Sequence-to-Sequence Models,”
C.-C. Chiu, T. N. Sainath, Y. Wu, et al., · 2018
Earlier work this paper cites.
“An analysis of incorporating an external language model into a sequence-to-sequence model,”
A. Kannan, Y. Wu, P. Nguyen, et al., · 2018
Earlier work this paper cites.
“Minimum Word Error Rate Training for Attention-based Sequence-to-sequence Models,”
R. Prabhavalkar, T. N. Sainath, Y. Wu, et al., · 2018
Earlier work this paper cites.
“Representation Learning with Contrastive Predictive Coding,”
A. Van Den Oord, Y. Li, and O. Vinyals, · 2018
Earlier work this paper cites.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,”
J. Devlin, M. Chang, K. Lee, and K. Toutanova, · 2018
Cited alongside, same era.
“Streaming End-to-end Speech Recognition For Mobile Devices,”
Y. He, T. N. Sainath, R. Prabhavalkar, et al., · 2019
Cited alongside, same era.
“Improving RNN transducer modeling for end-to-end speech recognition,”
J. Li, R. Zhao, H. Hu, and Y. Gong, · 2019
Cited alongside, same era.
“Cycle-Consistency Training for End-to-End Speech Recognition,”
T. Hori, R. Astudillo, T. Hayashi, et al., · 2019
Cited alongside, same era.
“On the Choice of Modeling Unit for Sequence-to-Sequence Speech Recognition,”
K. Irie, R. Prabhavalkar, A. Kannan, et al., · 2019
Cited alongside, same era.
“Recognizing Long-Form Speech Using Streaming End-to-End Models,”
A. Narayanan, R. Prabhavalkar, C.-C. Chiu, et al., · 2019
“An Efficient Streaming Non-Recurrent On-Device End-to-End Model with Improvements to Rare-Word Modeling,”
T. N. Sainath, Y. He, Narayanan, et al., · 2021
Later among the works it cites.
“Injecting Text in Self-Supervised Speech Pretraining,”
Z. Chen, Y. Zhang, A. Rosenberg, et al., · 2021
Later among the works it cites.
“SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training,”
A. Bapna, Y.-A. Chung, N. Wu, et al., · 2021
Later among the works it cites.
“SPLAT: Speech-Language Joint Pre-Training for Spoken Language Understanding,”
Y.-A. Chung, C. Zhu, and M. Zeng, · 2021
Later among the works it cites.
“Cascaded encoders for unifying streaming and non-streaming ASR,”
A. Narayanan, T. N. Sainath, R. Pang, et al., · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“wav2vec: Unsupervised Pre-training for Speech Recognition,”
S. Schneider, A. Baevski, R. Collobert, and M. Auli, · 2019
Cited alongside, same era.
M. Lewis, Y. Liu, N. Goyal, et al., · 2019
Cited alongside, same era.
“Joint Endpointing and Decoding with End-to-End Models,”
S. Chang, R. Prabhavalkar, Y. He, et al., · 2019
Cited alongside, same era.
“SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition,”
D. S. Park, W. Chan, Y. Zhang, et al., · 2019
Cited alongside, same era.
“Phoebe: Pronunciation-aware Contextualization for End-to-end Speech Recognition,”
A. Bruguier, R. Prabhavalkar, G. Pundak, and T. N. Sainath, · 2019
Cited alongside, same era.
“On the Comparison of Popular End-to-End Models for Large Scale Speech Recognition,”
J. Li, Y. Wu, Y. Gaur, et al., · 2020
Cited alongside, same era.
“HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,”
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, et al., · 2021
Later among the works it cites.
“w2v-bert: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training,”
Y.-A. Chung, Y. Zhang, W. Han, et al., · 2021
Later among the works it cites.
“Joint Masked CPC And CTC Training For ASR,”
C. Talnikar, T. Likhomanenko, R. Collobert, and G. Synnaeve, · 2021
Later among the works it cites.
“A Better and Faster End-to-End Model for Streaming ASR,”
B. Li, A. Gulati, J. Yu, et al., · 2021
Later among the works it cites.
“Tied & Reduced RNN-T Decoder,”
R. Botros and T. N. Sainath, · 2021
Later among the works it cites.
“FastEmit: Low-latency Streaming ASR with Sequence-level Emission Regularization,”
J. Yu, C.-C. Chiu, B. Li, et al., · 2021
Later among the works it cites.
“USTED: Improving ASR with a Unified Speech and Text Encoder-Decoder,”
B. Yusuf, A. Gandhe, and A. Sokolov, · 2022
Closest in time.
“mSLAM: Massively Multilingual Joint Pre-Training for Speech and Text,”
A. Bapna, C. Cherry, Y. Zhang, et al., · 2022
Closest in time.
“Unified Speech-Text Pre-training for Speech Translation and Recognition,”
Y. Tang, H. Gong, N. Dong, et al., · 2022
Closest in time.
“Towards Reducing the Need for Speech Training Data to Build Spoken Language Understanding Systems,”
S. Thomas, H. J. Kuo, B. Kingsbury, and G. Saon, · 2022
Closest in time.
“MAESTRO: Matched Speech Text Representations through Modality Matching,”
Z. Chen, Y. Zhang, A. Rosenberg, et al., · 2022
Closest in time.
“SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing,”
J. Ao, R. Wang, L. Zhou, et al., · 2022
Closest in time.
“Improving the Latency and Quality of Cascaded Encoder,”
T. N. Sainath, Y. He, A. Narayanan, et al., · 2022
Closest in time.
“Joint Unsupervised and Supervised Training for Multilingual ASR,”
J. Bai, B. Li, Y. Zhang, et al., · 2022
Closest in time.
“Self-Supervised Learning with Random-Projection Quantizer for Speech Recognition,”
C.-C. Chiu, J. Qin, Y. Zhang, et al., · 2022
Closest in time.
“Knowledge Transfer from Large-scale Pretrained Language Models to End-to-end Speech Recognizers,”
Y. Kubo, S. Karita, and M. Bacchiani, · 2022
Closest in time.
“Tts4pretrain 2.0: Advancing the use of Text and Speech in ASR Pretraining with Consistency and Contrastive Losses,”
Z. Chen, Y. Zhang, A. Rosenberg, B. Ramabhadran, P. Moreno, and G. Wang, · 2022
Closest in time.
“Pseudo Label Is Better Than Human Label,”
D. Hwang, K. Sim, Z. Huo, and T. Strohman, · 2022
Closest in time.