Fetching the paper…
Reading the bibliography…
In this paper, we propose an efficient and accurate streaming speech recognition model based on the FastConformer architecture.
“The ami meeting corpus,”
Wessel Kraaij, Thomas Hain, Mike Lincoln, and Wilfried Post, · 2005
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Ted-lium: An automatic speech recognition dedicated corpus,”
Anthony Rousseau, Paul Deléglise, and Yannick Estève, · 2012
Earlier work this paper cites.
“Batch normalization: Accelerating deep network training by reducing internal covariate shift,”
Sergey Ioffe and Christian Szegedy, · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Layer normalization,”
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton, · 2016
Earlier work this paper cites.
“Decoupled weight decay regularization,”
Ilya Loshchilov and Frank Hutter, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,”
Taku Kudo and John Richardson, · 2018
Earlier work this paper cites.
“Mixed precision training,”
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al., · 2018
Earlier work this paper cites.
“Streaming end-to-end speech recognition for mobile devices,”
Yanzhang He, Tara N Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Yonghui Wu, Ruoming Pang, et al., · 2019
Earlier work this paper cites.
“Nemo: a toolkit for building ai applications using neural modules,”
Oleksii Kuchaiev, Jason Li, Huyen Nguyen, Oleksii Hrinchuk, Ryan Leary, Boris Ginsburg, Samuel Kriman, Stanislav Beliaev, Vitaly Lavrukhin, Jack Cook, et al., · 2019
Cited alongside, same era.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al., · 2020
Cited alongside, same era.
“Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss,”
Qian Zhang, Han Lu, Hasim Sak, Anshuman Tripathi, Erik McDermott, Stephen Koo, and Shankar Kumar, · 2020
Cited alongside, same era.
“Streaming automatic speech recognition with the transformer model,”
Niko Moritz, Takaaki Hori, and Jonathan Le, · 2020
Cited alongside, same era.
“Streaming transformer-based acoustic models using self-attention with augmented memory,”
Chunyang Wu, Yongqiang Wang, Yangyang Shi, Ching-Feng Yeh, and Frank Zhang, · 2020
“Dual-mode ASR: Unify and improve streaming asr with full-context modeling,”
Jiahui Yu, Wei Han, Anmol Gulati, Chung-Cheng Chiu, Bo Li, Tara N Sainath, Yonghui Wu, and Ruoming Pang, · 2021
Later among the works it cites.
“Wenet: Production oriented streaming and non-streaming end-to-end speech recognition toolkit,”
Zhuoyuan Yao, Di Wu, Xiong Wang, Binbin Zhang, Fan Yu, Chao Yang, Zhendong Peng, Xiaoyu Chen, Lei Xie, and Xin Lei, · 2021
Later among the works it cites.
“Multi-mode transformer transducer with stochastic future context,”
Kwangyoun Kim, Felix Wu, Prashant Sridhar, Kyu J. Han, and Shinji Watanabe, · 2021
Later among the works it cites.
“Spgispeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition,”
Patrick K O’Neill, Vitaly Lavrukhin, Somshubra Majumdar, Vahid Noroozi, Yuekai Zhang, Oleksii Kuchaiev, Jagadeesh Balam, Yuliya Dovzhenko, Keenan Freyberg, Michael D Shulman, et al., · 2021
Later among the works it cites.
“Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Enhancing monotonic multihead attention for streaming asr,”
Hirofumi Inaguma, Masato Mimura, and Tatsuya Kawahara, · 2020
Cited alongside, same era.
“Low latency end-to-end streaming speech recognition with a scout network,”
Chengyi Wang, Yu Wu, Shujie Liu, Jinyu Li, Liang Lu, Guoli Ye, and Ming Zhou, · 2020
Cited alongside, same era.
“Synchronous transformers for end-to-end speech recognition,”
Zhengkun Tian, Jiangyan Yi, Ye Bai, Jianhua Tao, Shuai Zhang, and Zhengqi Wen, · 2020
Cited alongside, same era.
“Common voice: A massively-multilingual speech corpus,”
Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M Tyers, and Gregor Weber, · 2020
Cited alongside, same era.
“Developing real-time streaming transformer transducer for speech recognition on large-scale dataset,”
Xie Chen, Yu Wu, Zhenghao Wang, Shujie Liu, and Jinyu Li, · 2021
Cited alongside, same era.
“Streaming transformer asr with blockwise synchronous beam search,”
Emiru Tsunoo, Yosuke Kashiwagi, and Shinji Watanabe, · 2021
Cited alongside, same era.
“A better and faster end-to-end model for streaming asr,”
Bo Li, Anmol Gulati, Jiahui Yu, Tara N Sainath, Chung-Cheng Chiu, Arun Narayanan, Shuo-Yiin Chang, Ruoming Pang, Yanzhang He, James Qin, et al., · 2021
Cited alongside, same era.
Guoguo Chen, Shuzhou Chai, Guanbo Wang, Jiayu Du, Wei Qiang Zhang, Chao Weng, Dan Su, Daniel Povey, Jan Trmal, Junbo Zhang, et al., · 2021
Later among the works it cites.
“Voxpopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,”
Changhan Wang, Morgane Rivière, Ann Lee, Anne Wu, Chaitanya Talnikar, Daniel Haziza, Mary Williamson, Juan Pino, and Emmanuel Dupoux, · 2021
Later among the works it cites.
“Emformer: Efficient memory transformer based acoustic model for low latency streaming speech recognition,”
Yangyang Shi, Yongqiang Wang, Chunyang Wu, Ching-Feng Yeh, Julian Chan, Frank Zhang, Duc Le, and Mike Seltzer, · 2021
Later among the works it cites.
“Fastemit: Low-latency streaming asr with sequence-level emission regularization,”
Jiahui Yu, Chung-Cheng Chiu, Bo Li, Shuo-yiin Chang, Tara N Sainath, Yanzhang He, Arun Narayanan, Wei Han, Anmol Gulati, Yonghui Wu, et al., · 2021
Later among the works it cites.
“Earnings-22: A practical benchmark for accents in the wild,”
Miguel Del Rio, Peter Ha, Quinten McNamara, Corey Miller, and Shipra Chandra, · 2022
Later among the works it cites.
“Fast conformer with linearly scalable attention for efficient speech recognition,”
Dima Rekesh, Nithin Rao Koluguri, Samuel Kriman, Somshubra Majumdar, Vahid Noroozi, He Huang, Oleksii Hrinchuk, Krishna Puvvada, Ankur Kumar, Jagadeesh Balam, and Boris Ginsburg, · 2023
Closest in time.
“FastConformer Hybrid Large Streaming Multi (en-US),” https://huggingface.co/nvidia/stt_en_fastconformer_hybrid_large_streaming_multi
NVIDIA-NeMo, · 2023
Closest in time.