Fetching the paper…
Reading the bibliography…
In this paper, we introduce Libriheavy, a large-scale ASR corpus consisting of 50,000 hours of read English speech derived from LibriVox.
“Binary codes capable of correcting deletions, insertions, and reversals,”
Vladimir I Levenshtein et al., · 1966
Earlier work this paper cites.
“The design for the Wall Street Journal-based CSR corpus,”
Douglas B. Paul and Janet M. Baker, · 1992
Earlier work this paper cites.
“Switchboard-1 release 2 ldc97s62.,” 1993
Godfrey, John J., and Edward Holliman., · 1993
Earlier work this paper cites.
“Fisher english training speech part 1 transcripts ldc2004t19.,” 2004
Cieri, Christopher, et al., · 2004
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Linear work suffix array construction,”
Juha Kärkkäinen, Peter Sanders, and Stefan Burkhardt, · 2006
Earlier work this paper cites.
“Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition,”
George E Dahl, Dong Yu, et al., · 2011
Earlier work this paper cites.
“Speech recognition with deep recurrent neural networks,”
A. Mohamed A. Graves and G. Hinton, · 2013
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
D. Kingma and J. Ba, · 2014
Earlier work this paper cites.
“Librispeech: an ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, et al., · 2015
Earlier work this paper cites.
“Audio augmentation for speech recognition,”
Tom Ko, Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur, · 2015
Cited alongside, same era.
“Listen, attend and spell,”
William Chan, Navdeep Jaitly, et al., · 2016
Cited alongside, same era.
“Neural machine translation of rare words with subword units,”
Rico Sennrich, Barry Haddow, and Alexandra Birch, · 2016
Cited alongside, same era.
“Hybrid CTC/attention architecture for end-to-end speech recognition,”
Shinji Watanabe et al., · 2017
Cited alongside, same era.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, et al., · 2017
Cited alongside, same era.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, et al., · 2019
Cited alongside, same era.
“RNN-transducer with stateless prediction network,”
Mohammadreza Ghodsi, Xiaofeng Liu, James Apfel, et al., · 2020
Later among the works it cites.
“Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio,”
Guoguo Chen, Shuzhou Chai, Guanbo Wang, Jiayu Du, Daniel Povey, et al., · 2021
Later among the works it cites.
“The people’s speech: A large-scale diverse english speech recognition dataset for commercial usage,” 2021
Daniel Galvez, Greg Diamos, , et al., · 2021
Later among the works it cites.
“Attentive contextual carryover for multi-turn end-to-end spoken language understanding,”
Kai Wei, Thanh Tran, Feng-Ju Chang, et al., · 2021
Later among the works it cites.
“Lhotse: a speech data representation library for the modern deep learning ecosystem,” 2021
Piotr Żelasko, Daniel Povey, Jan ”Yenda” Trmal, and Sanjeev Khudanpur, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“MLS: A large-scale multilingual dataset for speech research,”
Vineel Pratap, Qiantong Xu, Anuroop Sriram, et al., · 2020
Cited alongside, same era.
“Libri-light: A benchmark for asr with limited or no supervision,”
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P. E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, et al., · 2020
Cited alongside, same era.
“Conformer: Convolution-augmented Transformer for Speech Recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, et al., · 2020
Cited alongside, same era.
“Wenet: Production oriented streaming and non-streaming end-to-end speech recognition toolkit,”
Zhuoyuan Yao, Di Wu, Xiong Wang, Binbin Zhang, et al., · 2021
Later among the works it cites.
“Context-aware end-to-end ASR using self-attentive embedding and tensor fusion,”
Shuo-Yiin Chang, Chao Zhang, Tara N Sainath, Bo Li, and Trevor Strohman, · 2023
Closest in time.
“Zipformer: A faster and better encoder for automatic speech recognition,”
Zengwei Yao, Liyong Guo, Xiaoyu Yang, et al., · 2023
Closest in time.
“Fast and parallel decoding for transducer,”
Wei Kang, Liyong Guo, Fangjun Kuang, Long Lin, Mingshuang Luo, Zengwei Yao, Xiaoyu Yang, Piotr Żelasko, and Daniel Povey, · 2023
Closest in time.