Fetching the paper…
Reading the bibliography…
Language identification greatly impacts the success of downstream tasks such as automatic speech recognition.
“Comparison of four approaches to automatic language identification of telephone speech,”
Marc A Zissman, · 1996
Earlier work this paper cites.
“A study on multilingual acoustic modeling for large vocabulary asr,”
Hui Lin, Li Deng, Dong Yu, Yi-fan Gong, Alex Acero, and Chin-Hui Lee, · 2009
Earlier work this paper cites.
“Multilingual acoustic modeling for speech recognition based on subspace gaussian mixture models,”
Lukáš Burget, Petr Schwarz, et al., · 2010
Earlier work this paper cites.
“Product quantization for nearest neighbor search,”
Herve Jegou, Matthijs Douze, and Cordelia Schmid, · 2010
Earlier work this paper cites.
“Current trends in multilingual speech processing,”
Hervé Bourlard, John Dines, et al., · 2011
Earlier work this paper cites.
“Multilingual acoustic models using distributed deep neural networks,”
Georg Heigold, Vincent Vanhoucke, Alan Senior, Patrick Nguyen, Marc’Aurelio Ranzato, Matthieu Devin, and Jeffrey Dean, · 2013
Earlier work this paper cites.
“Automatic language identification using deep neural networks,”
Ignacio Lopez-Moreno, Javier Gonzalez-Dominguez, et al., · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Layer normalization,”
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton, · 2016
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, et al., · 2017
Earlier work this paper cites.
“Categorical reparameterization with gumbel-softmax,”
Eric Jang, Shixiang Gu, and Ben Poole, · 2017
Earlier work this paper cites.
“A structured self-attentive sentence embedding,”
Zhouhan Lin, Minwei Feng, et al., · 2017
Earlier work this paper cites.
“Deep neural network embeddings for text-independent speaker verification.,”
David Snyder, Garcia-Romero, et al., · 2017
Cited alongside, same era.
“Speech-transformer: A no-recurrence sequence-to-sequence model for speech recognition,”
Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Cited alongside, same era.
“Speech2vec: A sequence-to-sequence framework for learning word embeddings from speech,”
Yu-An Chung and James Glass, · 2018
Cited alongside, same era.
“Multilingual sequence-to-sequence speech recognition: architecture, transfer learning, and language modeling,”
Jaejin Cho, Murali Karthick Baskar, et al., · 2018
Cited alongside, same era.
“Multilingual speech recognition with a single end-to-end model,”
Shubham Toshniwal, Tara N Sainath, Ron J Weiss, Bo Li, Pedro Moreno, Eugene Weinstein, and Kanishka Rao, · 2018
Cited alongside, same era.
“Transformer-based acoustic modeling for hybrid speech recognition,”
Yongqiang Wang, Abdelrahman Mohamed, Due Le, Chunxi Liu, Alex Xiao, Jay Mahadeokar, Hongzhao Huang, Andros Tjandra, Xiaohui Zhang, Frank Zhang, and et al., · 2020
Later among the works it cites.
“End-to-end ASR: from Supervised to Semi-Supervised Learning with Modern Architectures,”
Gabriel Synnaeve, Qiantong Xu, et al., · 2020
Later among the works it cites.
“Iterative pseudo-labeling for speech recognition,”
Qiantong Xu, Tatiana Likhomanenko, et al., · 2020
Later among the works it cites.
“Improved noisy student training for automatic speech recognition,”
Daniel S. Park, Yu Zhang, et al., · 2020
Later among the works it cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Later among the works it cites.
“Massively multilingual asr: 50 languages, 1 model, 1 billion parameters,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aäron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Cited alongside, same era.
“Large-scale multilingual speech recognition with a streaming end-to-end model,”
Anjuli Kannan, Arindrima Datta, et al., · 2019
Cited alongside, same era.
“Bytes are all you need: End-to-end multilingual speech recognition and synthesis with bytes,”
Bo Li, Yu Zhang, Tara Sainath, Yonghui Wu, and William Chan, · 2019
Cited alongside, same era.
“wav2vec: Unsupervised pre-training for speech recognition,”
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli, · 2019
Cited alongside, same era.
“vq-wav2vec: Self-supervised learning of discrete speech representations,”
Alexei Baevski, Steffen Schneider, and Michael Auli, · 2019
Cited alongside, same era.
“BERT: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Cited alongside, same era.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, et al., · 2020
Cited alongside, same era.
Vineel Pratap, Anuroop Sriram, et al., · 2020
Later among the works it cites.
“Exploring wav2vec 2.0 on speaker verification and language identification,”
Zhiyun Fan, Meng Li, Shiyu Zhou, and Bo Xu, · 2020
Later among the works it cites.
“Unsupervised cross-lingual representation learning for speech recognition,”
Alexis Conneau, Alexei Baevski, Ronan Collobert, Abdelrahman Mohamed, and Michael Auli, · 2020
Later among the works it cites.
“Pushing the limits of semi-supervised learning for automatic speech recognition,”
Yu Zhang, James Qin, et al., · 2020
Later among the works it cites.
“On layer normalization in the transformer architecture,”
Ruibin Xiong and Yunchang et al Yang, · 2020
Later among the works it cites.
“Robust wav2vec 2.0: Analyzing domain shift in self-supervised pre-training,”
Wei-Ning Hsu, Anuroop Sriram, et al., · 2021
Closest in time.