Fetching the paper…
Reading the bibliography…
We introduce the Universal Speech Model (USM), a single large model that performs automatic speech recognition (ASR) across 100+ languages.
J.-L. Gauvain, L. F. Lamel, G. Adda, and M. Adda-Decker, “The limsi continuous speech dictation system: evaluation on the arpa wall street journal task,” in
1994
Earlier work this paper cites.
G. Zavaliagkos and T. Colthurst, “Utilizing untranscribed training data to improve performance,” in
1998
Earlier work this paper cites.
F. Kubala, J. Davenport, H. Jin, D. Liu, T. Leek, S. Matsoukas, D. Miller, L. Nguyen, F. Richardson, R. Schwartz
1998
Earlier work this paper cites.
S. Chen, M. Gales, P. Gopalakrishnan, R. Gopinath, H. Printz, D. Kanevsky, P. Olsen, and L. Polymenakos, “Ibm’s lvcsr system for transcription of broadcast news used in the 1997 hub4 english evaluation,” in
1998
Earlier work this paper cites.
L. Lamel, J. luc Gauvain, and G. Adda, “Lightly supervised acoustic model training,” in
2000
Earlier work this paper cites.
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, W. Kraaij, M. Kronenthal
2005
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in
2006
Earlier work this paper cites.
2006
Earlier work this paper cites.
S. Novotney and R. Schwartz, “Analysis of low-resource acoustic model self-training,” in
2009
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,”
2012
Earlier work this paper cites.
M. Schuster and K. Nakajima, “Japanese and korean voice search,” in
2012
Earlier work this paper cites.
A. Rousseau, P. Deléglise, and Y. Esteve, “Ted-lium: an automatic speech recognition dedicated corpus.” in
2012
Earlier work this paper cites.
S. Thomas, M. L. Seltzer, K. Church, and H. Hermansky, “Deep neural network features and semi-supervised training for low resource speech recognition,” in
2013
Earlier work this paper cites.
M. J. F. Gales, K. Knill, A. Ragni, and S. P. Rath, “Speech recognition and keyword spotting for low-resource languages: Babel project research at cued,” in
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in
2016
Earlier work this paper cites.
2018
Earlier work this paper cites.
W.-N. Hsu and J. Glass, “Extracting domain invariant features by unsupervised learning for robust automatic speech recognition,” in
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. v. d. Oord, Y. Li, and O. Vinyals, “Representation learning with contrastive predictive coding,”
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
F. Hernandez, V. Nguyen, S. Ghannay, N. Tomashenko, and Y. Esteve, “Ted-lium 3: twice as much data and corpus repartition for experiments on speaker adaptation,” in
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Chorowski, R. J. Weiss, S. Bengio, and A. van den Oord, “Unsupervised speech representation learning using wavenet autoencoders,”
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
B. Li, T. N. Sainath, R. Pang, and Z. Wu, “Semi-supervised training for end-to-end models via weak distillation,” in
2019
Earlier work this paper cites.
G. Synnaeve, Q. Xu, J. Kahn, T. Likhomanenko, E. Grave, V. Pratap, A. Sriram, V. Liptchinsky, and R. Collobert, “End-to-end asr: from supervised to semi-supervised learning with modern architectures,” in
2019
Earlier work this paper cites.
S. H. K. Parthasarathi and N. Strom, “Lessons from building acoustic models with a million hours of speech,” in
2019
Earlier work this paper cites.
C.-C. Chiu, W. Han, Y. Zhang, R. Pang, S. Kishchenko, P. Nguyen, A. Narayanan, H. Liao, S. Zhang, A. Kannan
2019
Cited alongside, same era.
2019
Cited alongside, same era.
E. Tsunoo, Y. Kashiwagi, T. Kumakura, and S. Watanabe, “Transformer asr with contextual block processing,” in
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2021
Later among the works it cites.
B. Ramabhadran, K. Audhkhasi, P. J. M. Mengibar, and T. Chen, “Mixture model attention: Flexible streaming and non-streaming automatic speech recognition,” in
2021
Later among the works it cites.
X. Chen, Y. Wu, Z. Wang, S. Liu, and J. Li, “Developing real-time streaming transformer transducer for speech recognition on large-scale dataset,” in
2021
Later among the works it cites.
Y. Shi, Y. Wang, C. Wu, C.-F. Yeh, J. Chan, F. Zhang, D. Le, and M. Seltzer, “Emformer: Efficient memory transformer based acoustic model for low latency streaming speech recognition,” in
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2020
Cited alongside, same era.
Q. Xie, M.-T. Luong, E. Hovy, and Q. V. Le, “Self-training with noisy student improves imagenet classification,” in
2020
Cited alongside, same era.
2020
Cited alongside, same era.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu
2020
Cited alongside, same era.
S. Ling, Y. Liu, J. Salazar, and K. Kirchhoff, “Deep contextualized acoustic representations for semi-supervised speech recognition,” in
2020
Cited alongside, same era.
M. Riviere, A. Joulin, P.-E. Mazaré, and E. Dupoux, “Unsupervised pretraining transfers well across languages,” in
2020
Cited alongside, same era.
K. Tomanek, V. Zayats, D. Padfield, K. Vaillancourt, and F. Biadsy, “Residual adapters for parameter-efficient asr adaptation to atypical and accented speech,” in
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
N. P. Jouppi, D. H. Yoon, M. Ashcraft, M. Gottscho, T. B. Jablin, G. Kurian, J. Laudon, S. Li, P. Ma, X. Ma
2021
Later among the works it cites.
B. Li, A. Gulati, J. Yu, T. N. Sainath, C.-C. Chiu, A. Narayanan, S.-Y. Chang, R. Pang, Y. He, J. Qin
2021
Later among the works it cites.
2022
Later among the works it cites.
Y. Zhang, D. S. Park, W. Han, J. Qin, A. Gulati, J. Shor, A. Jansen, Y. Xu, Y. Huang, S. Wang, Z. Zhou, B. Li, M. Ma, W. Chan, J. Yu, Y. Wang, L. Cao, K. C. Sim, B. Ramabhadran, T. N. Sainath, F. Beaufays, Z. Chen, Q. V. Le, C.-C. Chiu, R. Pang, and Y. Wu, “Bigssl: Exploring the frontier of large-scale semi-supervised learning for automatic speech recognition,”
2022
Later among the works it cites.
2022
Later among the works it cites.
C.-C. Chiu, J. Qin, Y. Zhang, J. Yu, and Y. Wu, “Self-supervised learning with random-projection quantizer for speech recognition,” in
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
J. He, C. Zhou, X. Ma, T. Berg-Kirkpatrick, and G. Neubig, “Towards a unified view of parameter-efficient transfer learning,” in
2022
Later among the works it cites.
2022
Later among the works it cites.
S. Thomas, B. Kingsbury, G. Saon, and H.-K. J. Kuo, “Integrating text inputs for training and adapting rnn transducer asr models,” in
2022
Later among the works it cites.
Y. Cheng, Y. Zhang, M. Johnson, W. Macherey, and A. Bapna, “Mu
2022
Later among the works it cites.
Z.-H. Zhang, L. Zhou, J. Ao, S. Liu, L. Dai, J. Li, and F. Wei, “Speechut: Bridging speech and text with hidden-unit for encoder-decoder based speech-text pre-training,” in
2022
Later among the works it cites.
2022
Later among the works it cites.
S. Khurana, A. Laurent, and J. R. Glass, “Samu-xlsr: Semantically-aligned multimodal utterance-level cross-lingual speech representation,”
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
F. Biadsy, Y. Chen, X. Zhang, O. Rybakov, A. Rosenberg, and P. J. Moreno, “A scalable model specialization framework for training and inference using submodels and its application to speech model personalization,” in
2022
Later among the works it cites.
2022
Later among the works it cites.
T. N. Sainath, R. Prabhavalkar, A. Bapna, Y. Zhang, Z. Huo, Z. Chen, B. Li, W. Wang, and T. Strohman, “Joist: A joint speech and text streaming model for asr,” in
2023
Closest in time.
Z. Meng, W. Wang, R. Prabhavalkar, T. N. Sainath, T. Chen, E. Variani, Y. Zhang, B. Li, A. Rosenberg, and B. Ramabhadran, “Jeit: Joint end-to-end model and internal language model training for speech recognition,” in
2023
Closest in time.
Z. Meng, T. Chen, R. Prabhavalkar, Y. Zhang, G. Wang, K. Audhkhasi, J. Emond, T. Strohman, B. Ramabhadran, W. R. Huang
2023
Closest in time.