Fetching the paper…
Reading the bibliography…
This paper presents an overview and evaluation of some of the end-to-end ASR models on long-form audios.
“The design for the Wall Street Journal based CSR corpus,”
D. B. Paul and J. M. Baker, · 1992
Earlier work this paper cites.
“Godfrey, john and holliman, edward. 1997. switchboard-l release 2. philadelphia, pa: Linguis,”
Wiltrud Mihatsch, · 1997
Earlier work this paper cites.
“Fisher english training speech part 1 transcripts,”
Christopher Cieri, David Graff, Owen Kimball, Dave Miller, and Kevin Walker, · 2004
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Librispeech: an ASR corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“Wav2letter: an end-to-end convnet-based speech recognition system,”
Ronan Collobert, Christian Puhrsch, and Gabriel Synnaeve, · 2016
Earlier work this paper cites.
“Searching for activation functions,”
Prajit Ramachandran, Barret Zoph, and Quoc V. Le, · 2017
Earlier work this paper cites.
“Exploring architectures, data and units for streaming end-to-end speech recognition with RNN-Transducer,”
Kanishka Rao, Haşim Sak, and Rohit Prabhavalkar, · 2017
Earlier work this paper cites.
Accessed: 2018-04-06
“Mozilla: A journey to less than 10% word error rate,” https://hacks.mozilla.org/2017/11/a-journey-to-10-word-error-rate/ · 2018
Earlier work this paper cites.
“TED-LIUM 3: Twice as much data and corpus repartition for experiments on speaker adaptation,”
François Hernandez, Vincent Nguyen, Sahar Ghannay, Natalia Tomashenko, and Yannick Esteve, · 2018
Earlier work this paper cites.
“A comparison of end-to-end models for long-form speech recognition,”
Chung-Cheng Chiu, Wei Han, Yu Zhang, Ruoming Pang, Sergey Kishchenko, Patrick Nguyen, Arun Narayanan, Hank Liao, Shuyuan Zhang, Anjuli Kannan, et al., · 2019
Earlier work this paper cites.
“Recognizing long-form speech using streaming end-to-end models,”
Arun Narayanan, Rohit Prabhavalkar, Chung-Cheng Chiu, David Rybach, Tara N. Sainath, and Trevor Strohman, · 2019
Earlier work this paper cites.
“Jasper: An End-to-End Convolutional Neural Acoustic Model,”
Jason Li, Vitaly Lavrukhin, Boris Ginsburg, Ryan Leary, Oleksii Kuchaiev, Jonathan M. Cohen, Huyen Nguyen, and Ravi Teja Gadde, · 2019
Cited alongside, same era.
“Building the singapore english national speech corpus,”
Jia Xin Koh, Aqilah Mislan, Kevin Khoo, Brian Ang, Wilson Ang, Charmaine Ng, and YY Tan, · 2019
Cited alongside, same era.
“Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92),”
Junichi Yamagishi, Christophe Veaux, Kirsten MacDonald, et al., · 2019
Cited alongside, same era.
“Conformer: Convolution-augmented Transformer for Speech Recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang, · 2020
Cited alongside, same era.
“QuartzNet: Deep automatic speech recognition with 1d time-channel separable convolutions,”
Samuel Kriman, Stanislav Beliaev, Boris Ginsburg, Jocelyn Huang, Oleksii Kuchaiev, Vitaly Lavrukhin, Ryan Leary, Jason Li, and Yang Zhang, · 2020
Cited alongside, same era.
“VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,”
Changhan Wang, Morgane Riviere, Ann Lee, Anne Wu, Chaitanya Talnikar, Daniel Haziza, Mary Williamson, Juan Pino, and Emmanuel Dupoux, · 2021
Later among the works it cites.
“Europarl-ASR: A Large Corpus of Parliamentary Debates for Streaming ASR Benchmarking and Speech Data Filtering/Verbatimization,”
Gonçal V. Garcés Díaz-Munío, Joan Albert Silvestre-Cerdà, Javier Jorge, Adrià Giménez, Javier Iranzo-Sánchez, Pau Baquero-Arnal, Nahuel Roselló, Alejandro Pérez-González de Martos, Jorge Civera, Albert Sanchis, and Alfons Juan, · 2021
Later among the works it cites.
Daniel Galvez, Greg Diamos, Juan Ciro, Juan Felipe Cerón, Keith Achorn, Anjali Gopi, David Kanter, Maximilian Lam, Mark Mazumder, and Vijay Janapa Reddi, · 2021
Later among the works it cites.
“Earnings-21: A practical benchmark for ASR in the wild,”
Miguel Del Rio, Natalie Delworth, Ryan Westerman, Michelle Huang, Nishchal Bhandari, Joseph Palakapilly, Quinten McNamara, Joshua Dong, Piotr Zelasko, and Miguel Jetté, · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“ContextNet: Improving convolutional neural networks for automatic speech recognition with global context,”
Wei Han, Zhengdong Zhang, Yu Zhang, Jiahui Yu, Chung-Cheng Chiu, James Qin, Anmol Gulati, Ruoming Pang, and Yonghui Wu, · 2020
Cited alongside, same era.
“MLS: A large-scale multilingual dataset for speech research,”
Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve, and Ronan Collobert, · 2020
Cited alongside, same era.
Somshubra Majumdar, Jagadeesh Balam, Oleksii Hrinchuk, Vitaly Lavrukhin, Vahid Noroozi, and Boris Ginsburg, · 2021
Cited alongside, same era.
“A better and faster end-to-end model for streaming ASR,”
Bo Li, Anmol Gulati, Jiahui Yu, Tara N Sainath, Chung-Cheng Chiu, Arun Narayanan, Shuo-Yiin Chang, Ruoming Pang, Yanzhang He, James Qin, et al., · 2021
Cited alongside, same era.
“Dual-mode ASR: Unify and improve streaming ASR with full-context modeling,”
Jiahui Yu, Wei Han, Anmol Gulati, Chung-Cheng Chiu, Bo Li, Tara N Sainath, Yonghui Wu, and Ruoming Pang, · 2021
Cited alongside, same era.
“Wenet: Production oriented streaming and non-streaming end-to-end speech recognition toolkit,”
Zhuoyuan Yao, Di Wu, Xiong Wang, Binbin Zhang, Fan Yu, Chao Yang, Zhendong Peng, Xiaoyu Chen, Lei Xie, and Xin Lei, · 2021
Cited alongside, same era.
“Advanced long-context end-to-end speech recognition using context-expanded transformers,” 2021
Takaaki Hori, Niko Moritz, Chiori Hori, and Jonathan Le Roux, · 2021
Cited alongside, same era.
Later among the works it cites.
“The corpus of regional african american language,” 2021
Charlie Farrington and Tyler Kendall, · 2021
Later among the works it cites.
“E2e segmenter: Joint segmenting and decoding for long-form asr,”
W. Ronny Huang, Shuo yiin Chang, David Rybach, Rohit Prabhavalkar, Tara N. Sainath, Cyril Allauzen, Cal Peyser, and Zhiyun Lu, · 2022
Later among the works it cites.
“Streaming / Buffered ASR,” https://github.com/NVIDIA/NeMo/tree/main/examples/asr/asr_chunked_inference
NVIDIA NeMo, · 2022
Later among the works it cites.
“speech-datasets,” 6 2022
revdotcom, · 2022
Later among the works it cites.
“E2e segmentation in a two-pass cascaded encoder asr model,”
W. Ronny Huang, Shuo-Yiin Chang, Tara N. Sainath, Yanzhang He, David Rybach, Robert David, Rohit Prabhavalkar, Cyril Allauzen, Cal Peyser, and Trevor D. Strohman, · 2023
Closest in time.
“Fast conformer with linearly scalable attention for efficient speech recognition,”
Dima Rekesh, Nithin Rao Koluguri, Samuel Kriman, Somshubra Majumdar, Vahid Noroozi, He Huang, Oleksii Hrinchuk, Krishna Puvvada, Ankur Kumar, Jagadeesh Balam, and Boris Ginsburg, · 2023
Closest in time.
“Context-aware end-to-end asr using self-attentive embedding and tensor fusion,”
Shuo-Yiin Chang, Chao Zhang, Tara N. Sainath, Bo Li, and Trevor Strohman, · 2023
Closest in time.
“Robust speech recognition via large-scale weak supervision,”
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever, · 2023
Closest in time.