Fetching the paper…
Reading the bibliography…
We summarize the results of a host of efforts using giant automatic speech recognition (ASR) models pre-trained using large, diverse unlabeled datasets containing approximately a million hours of audio.
H. Scudder, “Probability of error of some adaptive pattern-recognition machines,”
1965
Earlier work this paper cites.
D. Yarowsky, “Unsupervised word sense disambiguation rivaling supervised methods,” in
1995
Earlier work this paper cites.
G. Zavaliagkos and T. Colthurst, “Utilizing untranscribed training data to improve performance,” in
1998
Earlier work this paper cites.
L. Lamel, J. luc Gauvain, and G. Adda, “Lightly supervised acoustic model training,” in
2000
Earlier work this paper cites.
E. Riloff and J. Wiebe, “Learning extraction patterns for subjective expressions,” in
2003
Earlier work this paper cites.
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, W. Kraaij, M. Kronenthal
2005
Earlier work this paper cites.
F. Boller and J. Becker, “Dementiabank database guide,”
2005
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in
2006
Earlier work this paper cites.
2006
Earlier work this paper cites.
S. Novotney and R. Schwartz, “Analysis of low-resource acoustic model self-training,” in
2009
Earlier work this paper cites.
S. Haq, P. J. Jackson, and J. Edge, “Speaker-dependent audio-visual emotion recognition.” in
2009
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz
2011
Earlier work this paper cites.
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg
2011
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,”
2012
Earlier work this paper cites.
M. Schuster and K. Nakajima, “Japanese and korean voice search,” in
2012
Earlier work this paper cites.
A. Rousseau, P. Deléglise, and Y. Esteve, “Ted-lium: an automatic speech recognition dedicated corpus.” in
2012
Earlier work this paper cites.
S. Thomas, M. L. Seltzer, K. Church, and H. Hermansky, “Deep neural network features and semi-supervised training for low resource speech recognition,” in
2013
Earlier work this paper cites.
H. Liao, E. McDermott, and A. Senior, “Large scale deep neural network acoustic modeling with semi-supervised training data for youtube video transcription,” in
2013
Earlier work this paper cites.
H. Cao, D. G. Cooper, M. K. Keutmann, R. C. Gur, A. Nenkova, and R. Verma, “CREMA-D: Crowd-sourced emotional multimodal actors dataset,”
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
K. J. Piczak, “ESC: Dataset for Environmental Sound Classification,” in
2015
Earlier work this paper cites.
R. Zazo Candil, T. N. Sainath, G. Simko, and C. Parada, “Feature learning with raw-waveform cldnns for voice activity detection,” 2016
2016
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Earlier work this paper cites.
A. Nagrani, J. S. Chung, and A. Zisserman, “Voxceleb: a large-scale speaker identification dataset,”
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
W.-N. Hsu and J. Glass, “Extracting domain invariant features by unsupervised learning for robust automatic speech recognition,” in
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. v. d. Oord, Y. Li, and O. Vinyals, “Representation learning with contrastive predictive coding,”
2018
Earlier work this paper cites.
M. Ravanelli, D. Serdyuk, and Y. Bengio, “Twin regularization for online speech recognition,”
2018
Earlier work this paper cites.
R. Takashima, S. Li, and H. Kawai, “An investigation of a knowledge distillation method for ctc acoustic models,” in
2018
Earlier work this paper cites.
G. Kurata and K. Audhkhasi, “Improved knowledge distillation from bi-directional to uni-directional lstm ctc for end-to-end speech recognition,” in
2018
Cited alongside, same era.
A. Narayanan, A. Misra, K. C. Sim, G. Pundak, A. Tripathi, M. Elfeky, P. Haghani, T. Strohman, and M. Bacchiani, “Toward domain-invariant speech recognition via large scale training,” in
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
N. Shazeer and M. Stern, “Adafactor: Adaptive learning rates with sublinear memory cost,” in
2020
Later among the works it cites.
J. Kahn, M. Riviere, W. Zheng, E. Kharitonov, Q. Xu, P. Mazare, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, and et al., “Libri-light: A benchmark for asr with limited or no supervision,”
2020
Later among the works it cites.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu
2020
Later among the works it cites.
J. Shor, A. Jansen, R. Maor, O. Lang, O. Tuval, F. de Chaumont Quitry, M. Tagliasacchi, I. Shavitt, D. Emanuel, and Y. Haviv, “Towards Learning a Universal Non-Semantic Representation of Speech,” in
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
F. Hernandez, V. Nguyen, S. Ghannay, N. Tomashenko, and Y. Esteve, “Ted-lium 3: twice as much data and corpus repartition for experiments on speaker adaptation,” in
2018
Cited alongside, same era.
C. Boeddecker, J. Heitkaemper, J. Schmalenstroeer, L. Drude, J. Heymann, and R. Haeb-Umbach, “Front-end processing for the CHiME-5 dinner party scenario,” in
2018
Cited alongside, same era.
K. MacLean, “Voxforge,”
2018
Cited alongside, same era.
P. Warden, “Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition,”
2018
Cited alongside, same era.
A. Jansen, M. Plakal, R. Pandya, D. P. Ellis, S. Hershey, J. Liu, R. C. Moore, and R. A. Saurous, “Unsupervised learning of semantic audio representations,” in
2018
Cited alongside, same era.
2019
Cited alongside, same era.
J. Chorowski, R. J. Weiss, S. Bengio, and A. van den Oord, “Unsupervised speech representation learning using wavenet autoencoders,”
2019
Cited alongside, same era.
Q. Xie, M.-T. Luong, E. Hovy, and Q. V. Le, “Self-training with noisy student improves imagenet classification,” in
2020
Later among the works it cites.
D. S. Park, Y. Zhang, C.-C. Chiu, Y. Chen, B. Li, W. Chan, Q. V. Le, and Y. Wu, “Specaugment on large scale datasets,” in
2020
Later among the works it cites.
G. Kurata and G. Saon, “Knowledge distillation from offline to streaming rnn transducer for end-to-end speech recognition.” in
2020
Later among the works it cites.
J. Yu, W. Han, A. Gulati, C.-C. Chiu, B. Li, T. N. Sainath, Y. Wu, and R. Pang, “Universal asr: Unify and improve streaming asr with full-context modeling,”
2020
Later among the works it cites.
2020
Later among the works it cites.
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He, “Zero: Memory optimizations toward training trillion parameter models,” in
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
I. Medennikov, M. Korenevsky, T. Prisyach, Y. Khokhlov, M. Korenevskaya, I. Sorokin, T. Timofeeva, A. Mitrofanov, A. Andrusenko, I. Podluzhny
2020
Later among the works it cites.
Q. Zhang, H. Lu, H. Sak, A. Tripathi, E. McDermott, S. Koo, and S. Kumar, “Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss,” in
2020
Later among the works it cites.
B. W. Schuller, A. Batliner, C. Bergler, E.-M. Messner, A. Hamilton, S. Amiriparian, A. Baird, G. Rizos, M. Schmitt, L. Stappen, H. Baumeister, A. D. MacIntyre, and S. Hantke, “The INTERSPEECH 2020 Computational Paralinguistics Challenge: Elderly emotion, Breathing & Masks,” in
2020
Later among the works it cites.
J. Szep and S. Hariri, “Paralinguistic Classification of Mask Wearing by Image Classifiers and Fusion,” in
2020
Later among the works it cites.
M. Plakal and D. Ellis, “Yamnet,” Jan 2020. [Online]. Available:
2020
Later among the works it cites.
2021
Closest in time.
Z. Chen, A. Rosenberg, Y. Zhang, H. Zen, M. Ghodsi, Y. Huang, J. Emond, G. Wang, B. Ramabhadran, and P. J. M. Mengibar, “Semi-supervision in asr: Sequential mixmatch and factorized tts-based augmentation,” 2021
2021
Closest in time.
Q. Xu, A. Baevski, T. Likhomanenko, P. Tomasello, A. Conneau, R. Collobert, G. Synnaeve, and M. Auli, “Self-training and pre-training are complementary for speech recognition,” in
2021
Closest in time.
2021
Closest in time.
B. Li, A. Gulati, J. Yu, T. N. Sainath, C.-C. Chiu, A. Narayanan, S.-Y. Chang, R. Pang, Y. He, J. Qin
2021
Closest in time.
2021
Closest in time.
T. Doutre, W. Han, M. Ma, Z. Lu, C. Chiu, R. Pang, A. Narayanan, A. Misra, Y. Zhang, and L. Cao, “Improving streaming automatic speech recognition with non-streaming model distillation on unsupervised data,”
2021
Closest in time.
T. Doutre, W. Han, C. Chiu, R. Pang, O. Siohan, and L. Cao, “Bridging the gap between streaming and non-streaming ASR systems bydistilling ensembles of CTC and RNN-T models,”
2021
Closest in time.
2021
Closest in time.
C.-C. Chiu, A. Narayanan, W. Han, R. Prabhavalkar, Y. Zhang, N. Jaitly, R. Pang, T. N. Sainath, P. Nguyen, L. Cao
2021
Closest in time.
2021
Closest in time.
Z. Tüske, G. Saon, and B. Kingsbury, “On the limit of english conversational speech recognition,”
2021
Closest in time.
J. Peplinski, J. Shor, S. Joglekar, J. Garrison, and S. Patel, “FRILL: A Non-Semantic Speech Embedding for Mobile Devices,” in
2021
Closest in time.
D. Seo, H.-S. Oh, and Y. Jung, “Wav2kws: Transfer learning from speech representations for keyword spotting,”
2021
Closest in time.
S. Venugopalan, J. Shor, M. Plakal, J. Tobin, K. Tomanek, J. R. Green, and M. P. Brenner, “Comparing Supervised Models and Learned Speech Representations for Classifying Intelligibility of Disordered Speech on Selected Phrases,” in
2021
Closest in time.