Fetching the paper…
Reading the bibliography…
Expanding the language coverage of speech technology has the potential to improve access to information for many more people.
Ascii phonetic symbols for the world’s languages: Worldbet
J. L. Hieronymus · 1993
Earlier work this paper cites.
Computer-coding the ipa: a proposed extension of sampa, 1995
J. Wells · 1995
Earlier work this paper cites.
Unsupervised cross-lingual representation learning for speech recognition
A. Conneau, A. Baevski, R. Collobert, A. Mohamed, and M. Auli · 2006
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
A. Graves, S. Fernández, and F. Gomez · 2006
Earlier work this paper cites.
Massively multilingual asr: 50 languages, 1 model, 1 billion parameters
V. Pratap, A. Sriram, et al · 2007
Earlier work this paper cites.
Multilingual transliteration using feature based phonetic method
S.-Y. Yoon, K.-Y. Kim, and R. Sproat · 2007
Earlier work this paper cites.
A study on multilingual acoustic modeling for large vocabulary asr
H. Lin, L. Deng, D. Yu, Y.-f. Gong, A. Acero, and C.-H. Lee · 2009
Earlier work this paper cites.
Multilingual acoustic modeling for speech recognition based on subspace gaussian mixture models
L. Burget, P. Schwarz, et al · 2010
Earlier work this paper cites.
A python toolkit for universal transliteration
T. Qian, K. Hollingshead, S.-y. Yoon, K.-y. Kim, and R. Sproat · 2010
Earlier work this paper cites.
Current trends in multilingual speech processing
H. Bourlard, J. Dines, et al · 2011
Earlier work this paper cites.
KenLM: Faster and smaller language model queries
K. Heafield · 2011
Earlier work this paper cites.
Product quantization for nearest neighbor search
H. Jegou, M. Douze, and C. Schmid · 2011
Earlier work this paper cites.
The kaldi speech recognition toolkit
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely · 2011
Earlier work this paper cites.
Crowdmos: An approach for crowdsourcing mean opinion score studies
F. Ribeiro, D. Florêncio, C. Zhang, and M. Seltzer · 2011
Earlier work this paper cites.
Multilingual acoustic models using distributed deep neural networks
G. Heigold, V. Vanhoucke, A. Senior, P. Nguyen, M. Ranzato, M. Devin, and J. Dean · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
Speech recognition and keyword spotting for low-resource languages: Babel project research at cued
M. J. F. Gales, K. M. Knill, A. Ragni, and S. P. Rath · 2014
Earlier work this paper cites.
Pyin: A fundamental frequency estimator using probabilistic threshold distributions
M. Mauch and S. Dixon · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Earlier work this paper cites.
Listen, attend and spell
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals · 2015
Earlier work this paper cites.
A massively parallel corpus: the bible in 100 languages
C. Christodouloupoulos and M. Steedman · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Layer normalization
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Training deep nets with sublinear memory cost
T. Chen, B. Xu, C. Zhang, and C. Guestrin · 2016
Earlier work this paper cites.
Wav2letter: an end-to-end convnet-based speech recognition system
R. Collobert, C. Puhrsch, and G. Synnaeve · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
E. Jang, S. Gu, and B. Poole · 2016
Earlier work this paper cites.
Ethnologue: Languages of the world, nineteenth edition
M. P. Lewis, G. F. Simon, and C. D. Fennig · 2016
Earlier work this paper cites.
Density estimation using real NVP
L. Dinh, J. Sohl-Dickstein, and S. Bengio · 2017
Earlier work this paper cites.
The lj speech dataset
K. Ito and L. Johnson · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Multilingual sequence-to-sequence speech recognition: architecture, transfer learning, and language modeling
J. Cho, M. K. Baskar, et al · 2018
Earlier work this paper cites.
The challenge of realistic music generation: modelling raw audio at scale
S. Dieleman, A. van den Oord, and K. Simonyan · 2018
Cited alongside, same era.
Ina’s mirex 2018 music and speech detection system
D. Doukhan, E. Lechapt, M. Evrard, and J. Carrive · 2018
Cited alongside, same era.
Out-of-the-box universal Romanization tool uroman
U. Hermjakob, J. May, and K. Knight · 2018
Cited alongside, same era.
Multilingual speech recognition with a single end-to-end model
S. Toshniwal, T. N. Sainath, R. J. Weiss, B. Li, P. Moreno, E. Weinstein, and K. Rao · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
A. van den Oord, Y. Li, and O. Vinyals · 2018
Cited alongside, same era.
ESPnet: End-to-end speech processing toolkit
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. Enrique Yalta Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, and T. Ochiai · 2018
Fairscale: A general purpose modular pytorch library for high performance and large scale training
M. Baines, S. Bhosale, V. Caggiano, N. Goyal, S. Goyal, M. Ott, B. Lefaudeux, V. Liptchinsky, M. Rabbat, S. Sheiffer, A. Sridhar, and M. Xu · 2021
Later among the works it cites.
Global predictors of language endangerment and the future of linguistic diversity
L. Bromham, R. Dinnage, H. Skirgard, A. Ritchie, M. Cardillo, F. Meakins, S. Greenhill, and X. Hua · 2021
Later among the works it cites.
Speechstew: Simply mix all available speech recognition data to train one large neural network
W. Chan, D. Park, C. Lee, Y. Zhang, Q. Le, and M. Norouzi · 2021
Later among the works it cites.
Exploring wav2vec 2.0 on speaker verification and language identification
Z. Fan, M. Li, S. Zhou, and B. Xu · 2021
Later among the works it cites.
Multilingual byte2speech models for scalable low-resource speech synthesis
M. He, J. Yang, L. He, and F. K. Soong · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Cmu wilderness multilingual speech dataset
A. W. Black · 2019
Cited alongside, same era.
Cross-lingual language model pretraining
A. Conneau and G. Lample · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for nlp
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly · 2019
Cited alongside, same era.
Large-scale multilingual speech recognition with a streaming end-to-end model
A. Kannan, A. Datta, et al · 2019
Cited alongside, same era.
Bytes are all you need: End-to-end multilingual speech recognition and synthesis with bytes
B. Li, Y. Zhang, T. Sainath, Y. Wu, and W. Chan · 2019
Cited alongside, same era.
Later among the works it cites.
Robust wav2vec 2.0: Analyzing domain shift in self-supervised pre-training
W.-N. Hsu, A. Sriram, A. Baevski, T. Likhomanenko, Q. Xu, V. Pratap, J. Kahn, A. Lee, R. Collobert, G. Synnaeve, et al · 2021
Later among the works it cites.
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech
J. Kim, J. Kong, and J. Son · 2021
Later among the works it cites.
Scaling end-to-end models for large-scale multilingual asr
B. Li, R. Pang, T. N. Sainath, A. Gulati, Y. Zhang, J. Qin, P. Haghani, W. R. Huang, M. Ma, and J. Bai · 2021
Later among the works it cites.
Zero-infinity: Breaking the GPU memory wall for extreme scale deep learning
S. Rajbhandari, O. Ruwase, J. Rasley, S. Smith, and Y. He · 2021
Later among the works it cites.
SpeechBrain: A general-purpose speech toolkit, 2021
M. Ravanelli, T. Parcollet, P. Plantinga, A. Rouhe, S. Cornell, L. Lugosch, C. Subakan, N. Dawalatabad, A. Heba, J. Zhong, J.-C. Chou, S.-L. Yeh, S.-W. Fu, C.-F. Liao, E. Rastorgueva, F. Grondin, W. Aris, H. Na, Y. Gao, R. D. Mori, and Y. Bengio · 2021
Later among the works it cites.
A survey on neural speech synthesis
X. Tan, T. Qin, F. Soong, and T.-Y. Liu · 2021
Later among the works it cites.
VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation
C. Wang, M. Riviere, A. Lee, A. Wu, C. Talnikar, D. Haziza, M. Williamson, J. Pino, and E. Dupoux · 2021
Later among the works it cites.
Torchaudio: Building blocks for audio and speech processing
Y.-Y. Yang, M. Hira, Z. Ni, A. Chourdia, A. Astafurov, C. Chen, C.-F. Yeh, C. Puhrsch, D. Pollack, D. Genzel, D. Greenberg, E. Z. Yang, J. Lian, J. Mahadeokar, J. Hwang, J. Chen, P. Goldsborough, P. Roy, S. Narenthiran, S. Watanabe, S. Chintala, V. Quenneville-Bélair, and Y. Shi · 2021
Later among the works it cites.
XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale
A. Babu, C. Wang, A. Tjandra, K. Lakhotia, Q. Xu, N. Goyal, K. Singh, P. von Platen, Y. Saraf, J. Pino, A. Baevski, A. Conneau, and M. Auli · 2022
Later among the works it cites.
W-CTC: a connectionist temporal classification loss with wild cards
X. Cai, J. Yuan, Y. Bian, G. Xun, J. Huang, and K. Church · 2022
Later among the works it cites.
Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone
E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. Gölge, and M. A. Ponti · 2022
Later among the works it cites.
Maestro: Matched speech text representations through modality matching
Z. Chen, Y. Zhang, A. Rosenberg, B. Ramabhadran, P. Moreno, A. Bapna, and H. Zen · 2022
Later among the works it cites.
Fleurs: Few-shot learning evaluation of universal representations of speech
A. Conneau, M. Ma, S. Khanuja, Y. Zhang, V. Axelrod, S. Dalmia, J. Riesa, C. Rivera, and A. Bapna · 2022
Later among the works it cites.
Towards building asr systems for the next billion users
T. Javed, S. Doddapaneni, A. Raman, K. S. Bhogale, G. Ramesh, A. Kunchukuttan, P. Kumar, and M. M. Khapra · 2022
Later among the works it cites.
Ambernet: A compact end-to-end model for spoken language identification
F. Jia, N. R. Koluguri, J. Balam, and B. Ginsburg · 2022
Later among the works it cites.
Flashlight: Enabling innovation in tools for machine learning
J. D. Kahn, V. Pratap, T. Likhomanenko, Q. Xu, A. Hannun, J. Cai, P. Tomasello, A. Lee, E. Grave, G. Avidov, et al · 2022
Later among the works it cites.
Bloom library: Multimodal datasets in 300+ languages for a variety of downstream tasks
C. Leong, J. Nemecek, J. Mansdorfer, A. Filighera, A. Owodunni, and D. Whitenack · 2022
Later among the works it cites.
Asr2k: Speech recognition for around 2000 languages without audio
X. Li, F. Metze, D. R. Mortensen, A. W. Black, and S. Watanabe · 2022
Later among the works it cites.
Towards end-to-end unsupervised speech recognition
A. H. Liu, W.-N. Hsu, M. Auli, and A. Baevski · 2022
Later among the works it cites.
Pseudo-labeling for massively multilingual speech recognition
L. Lugosch, T. Likhomanenko, G. Synnaeve, and R. Collobert · 2022
Later among the works it cites.
Bibletts: a large, high-fidelity, multilingual, and uniquely african speech corpus
J. Meyer, D. I. Adelani, E. Casanova, A. Öktem, D. W. J. Weber, S. Kabongo, E. Salesky, I. Orife, C. Leong, P. Ogayo, C. Emezue, J. Mukiibi, S. Osei, A. Agbolo, V. Akinode, B. Opoku, S. Olanrewaju, J. Alabi, and S. Muhammad · 2022
Later among the works it cites.
No language left behind: Scaling human-centered machine translation, 2022
NLLB_Team, M. R. Costa-jussà, J. Cross, O. Çelebi, M. Elbayad, K. Heafield, K. Heffernan, E. Kalbassi, J. Lam, D. Licht, J. Maillard, A. Sun, S. Wang, G. Wenzek, A. Youngblood, B. Akula, L. Barrault, G. M. Gonzalez, P. Hansanti, J. Hoffman, S. Jarrett, K. R. Sadagopan, D. Rowe, S. Spruit, C. Tran, P. Andrews, N. F. Ayan, S. Bhosale, S. Edunov, A. Fan, C. Gao, V. Goswami, F. Guzmán, P. Koehn, A. Mourachko, C. Ropers, S. Saleem, H. Schwenk, and J. Wang · 2022
Later among the works it cites.
Star temporal classification: Sequence modeling with partially labeled data
V. Pratap, A. Hannun, G. Synnaeve, and R. Collobert · 2022
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever · 2022
Later among the works it cites.
Improved language identification through cross-lingual self-supervised learning
A. Tjandra, D. G. Choudhury, F. Zhang, K. Singh, A. Conneau, A. Baevski, A. Sela, Y. Saraf, and M. Auli · 2022
Later among the works it cites.
Improving massively multilingual asr with auxiliary ctc objectives
W. Chen, B. Yan, J. Shi, Y. Peng, S. Maiti, and S. Watanabe · 2023
Closest in time.