Fetching the paper…
Reading the bibliography…
Building ASR models across many languages is a challenging multi-task learning problem due to large variations and heavily unbalanced data.
The need for biases in learning generalizations
Tom M Mitchell, · 1980
Earlier work this paper cites.
“ASCII phonetic symbols for the world’s languages: Worldbet,”
James L Hieronymus, · 1993
Earlier work this paper cites.
“Computer-coding the IPA: a proposed extension of SAMPA,”
John C Wells, · 1995
Earlier work this paper cites.
“Multitask learning,”
Rich Caruana, · 1997
Earlier work this paper cites.
Handbook of the International Phonetic Association: A guide to the use of the International Phonetic Alphabet
International Phonetic Association, International Phonetic Association Staff, et al., · 1999
Earlier work this paper cites.
“ASR-articulatory speech recognition,”
Joe Frankel and Simon King, · 2001
Earlier work this paper cites.
“Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,”
Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, et al., · 2012
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Long Short-Term Memory Recurrent Neural Network Architectures for Large Scale Acoustic Modeling,”
H. Sak, A. Senior, and F. Beaufays, · 2014
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, · 2015
Earlier work this paper cites.
“Tensorflow: Large-scale machine learning on heterogeneous distributed systems,” Available online: http://download.tensorflow.org/paper/whitepaper2015.pdf, 2015
M. Abadi et al., · 2015
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
D. P. Kingma and J. Ba, · 2015
Earlier work this paper cites.
“Multilingual techniques for low resource automatic speech recognition,”
Ekapol Chuangsuwanich, · 2016
Earlier work this paper cites.
“Joint CTC-attention based end-to-end speech recognition using multi-task learning,”
S. Kim, T. Hori, and S. Watanabe, · 2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, · 2017
Earlier work this paper cites.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2018
Cited alongside, same era.
“Multi-Dialect Speech Recognition With A Single Sequence-To-Sequence Model,”
Bo Li, Tara N Sainath, Khe Chai Sim, Michiel Bacchiani, Eugene Weinstein, Patrick Nguyen, Zhifeng Chen, Yanghui Wu, and Kanishka Rao, · 2018
Cited alongside, same era.
“Multilingual end-to-end speech recognition with a single transformer on low-resource languages,”
Shiyu Zhou, Shuang Xu, and Bo Xu, · 2018
Cited alongside, same era.
“Adafactor: Adaptive learning rates with sublinear memory cost,”
Noam Shazeer and Mitchell Stern, · 2018
Cited alongside, same era.
“Streaming End-to-end Speech Recognition For Mobile Devices,”
Y. He, T. N. Sainath, R. Prabhavalkar, et al., · 2019
Cited alongside, same era.
“On the Comparison of Popular End-to-End Models for Large Scale Speech Recognition,”
J. Li, Y. Wu, Y. Gaur, et al., · 2020
Later among the works it cites.
“Transformer transducer: A streamable speech recognition model with transformer encoders and RNN-T loss,”
Qian Zhang, Han Lu, Hasim Sak, Anshuman Tripathi, Erik McDermott, Stephen Koo, and Shankar Kumar, · 2020
Later among the works it cites.
“Massively multilingual ASR: 50 Languages, 1 Model, 1 Billion Parameters,”
Vineel Pratap, Anuroop Sriram, Paden Tomasello, et al., · 2020
Later among the works it cites.
“Large-Scale End-to-End Multilingual Speech Recognition and Language Identification with Multi-Task Learning,”
Wenxin Hou, Yue Dong, Bairong Zhuang, Longfei Yang, Jiatong Shi, and Takahiro Shinozaki, · 2020
Later among the works it cites.
“Balancing training for multilingual neural machine translation,”
Xinyi Wang, Yulia Tsvetkov, and Graham Neubig, · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges,”
Naveen Arivazhagan, Ankur Bapna, Orhan Firat, Dmitry Lepikhin, Melvin Johnson, Maxim Krikun, Mia Xu Chen, Yuan Cao, George Foster, Colin Cherry, et al., · 2019
Cited alongside, same era.
“Large-Scale Multilingual Speech Recognition with a Streaming End-to-End Model,”
Anjuli Kannan, Arindrima Datta, Tara N Sainath, Eugene Weinstein, Bhuvana Ramabhadran, Yonghui Wu, Ankur Bapna, Zhifeng Chen, and Seungji Lee, · 2019
Cited alongside, same era.
“Bytes are All You Need: End-to-End Multilingual Speech Recognition and Synthesis with Bytes,”
Bo Li, Yu Zhang, Tara Sainath, Yonghui Wu, and William Chan, · 2019
Cited alongside, same era.
“Massively multilingual adversarial speech recognition,”
Oliver Adams, Matthew Wiesner, Shinji Watanabe, and David Yarowsky, · 2019
Cited alongside, same era.
“Lingvo: a modular and scalable framework for sequence-to-sequence modeling,”
J. Shen, P. Nguyen, Y. Wu, et al., · 2019
Cited alongside, same era.
“SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition,”
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le, · 2019
Cited alongside, same era.
“GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding,”
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen, · 2020
Cited alongside, same era.
Later among the works it cites.
Zirui Wang, Yulia Tsvetkov, Orhan Firat, and Yuan Cao, · 2020
Later among the works it cites.
“Knowledge distillation for multi-task learning,”
Wei-Hong Li and Hakan Bilen, · 2020
Later among the works it cites.
“Scaling laws for neural language models,”
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei, · 2020
Later among the works it cites.
“Conformer: Convolution-augmented Transformer for Speech Recognition,”
A. Gulati, J. Qin, C.-C. Chiu, et al., · 2020
Later among the works it cites.
“Contextnet: Improving convolutional neural networks for automatic speech recognition with global context,”
Wei Han, Zhengdong Zhang, Yu Zhang, Jiahui Yu, Chung-Cheng Chiu, James Qin, Anmol Gulati, Ruoming Pang, and Yonghui Wu, · 2020
Later among the works it cites.
“MLS: A Large-Scale Multilingual Dataset for Speech Research,”
Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve, and Ronan Collobert, · 2020
Later among the works it cites.
“Unsupervised cross-lingual representation learning for speech recognition,”
Alexis Conneau, Alexei Baevski, Ronan Collobert, Abdelrahman Mohamed, and Michael Auli, · 2020
Later among the works it cites.
“Advancing rnn transducer technology for speech recognition,”
George Saon, Zoltan Tueske, Daniel Bolanos, and Brian Kingsbury, · 2021
Closest in time.
“A Better and Faster End-to-End Model for Streaming ASR,”
B. Li, A. Gulati, J. Yu, et al., · 2021
Closest in time.