Fetching the paper…
Reading the bibliography…
Despite rapid progress in the recent past, current speech recognition systems still require labeled training data which limits this technology to a small fraction of the languages spoken around the globe.
A maximum likelihood approach to continuous speech recognition
L. R. Bahl, F. Jelinek, and R. L. Mercer · 1983
Earlier work this paper cites.
Cross-language speech perception: Evidence for perceptual reorganization during the first year of life
J. F. Werker and R. C. Tees · 1984
Earlier work this paper cites.
Maximum mutual information estimation of hidden markov model parameters for speech recognition
L. Bahl, P. Brown, P. De Souza, and R. Mercer · 1986
Earlier work this paper cites.
Maximum likelihood estimation for multivariate mixture observations of markov chains (corresp.)
B.-H. Juang, S. Levinson, and M. Sondhi · 1986
Earlier work this paper cites.
Clauses are perceptual units for young infants
K. Hirsh-Pasek, D. G. Kemler Nelson, P. W. Jusczyk, K. W. Cassidy, B. Druss, and L. Kennedy · 1987
Earlier work this paper cites.
The DARPA TIMIT Acoustic-Phonetic Continuous Speech Corpus CDROM
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, D. S. Pallett, and N. L. Dahlgren · 1993
Earlier work this paper cites.
Developmental changes in perception of nonnative vowel contrasts
L. Polka and J. F. Werker · 1994
Earlier work this paper cites.
Large vocabulary continuous speech recognition: A review
S. Young · 1996
Earlier work this paper cites.
Finite-state transducers in language and speech processing
M. Mohri · 1997
Earlier work this paper cites.
The beginnings of word segmentation in english-learning infants
P. W. Jusczyk, D. M. Houston, and M. Newsome · 1999
Earlier work this paper cites.
Word segmentation by 8-month-olds: When speech cues count more than statistics
E. K. Johnson and P. W. Jusczyk · 2001
Earlier work this paper cites.
Weighted finite-state transducers in speech recognition
M. Mohri, F. Pereira, and M. Riley · 2002
Earlier work this paper cites.
An amharic speech corpus for large vocabulary continuous speech recognition
S. T. Abate, W. Menzel, and B. Tafila · 2005
Earlier work this paper cites.
Discriminative training for large vocabulary speech recognition
D. Povey · 2005
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
A. Graves, S. Fernández, and F. Gomez · 2006
Earlier work this paper cites.
Unsupervised learning of acoustic sub-word units
B. Varadarajan, S. Khudanpur, and E. Dupoux · 2008
Earlier work this paper cites.
Unsupervised training of an hmm-based speech recognizer for topic classification
H. Gish, M. Siu, A. Chan, and W. Belfield · 2009
Earlier work this paper cites.
Unsupervised spoken keyword spotting via segmental dtw on gaussian posteriorgrams
Y. Zhang and J. R. Glass · 2009
Earlier work this paper cites.
Improving translation model by monolingual data
O. Bojar and A. Tamchyna · 2011
Earlier work this paper cites.
KenLM: Faster and smaller language model queries
K. Heafield · 2011
Earlier work this paper cites.
Product quantization for nearest neighbor search
H. Jegou, M. Douze, and C. Schmid · 2011
Earlier work this paper cites.
The kaldi speech recognition toolkit
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely · 2011
Earlier work this paper cites.
Applying convolutional neural networks concepts to hybrid nn-hmm model for speech recognition
O. Abdel-Hamid, A. Mohamed, H. Jiang, and G. Penn · 2012
Earlier work this paper cites.
Connectionist speech recognition: a hybrid approach , volume 247
H. A. Bourlard and N. Morgan · 2012
Earlier work this paper cites.
Developments of Swahili resources for an automatic speech recognition system
H. Gelas, L. Besacier, and F. Pellegrino · 2012
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
A. Graves · 2012
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, et al · 2012
Earlier work this paper cites.
A nonparametric bayesian approach to acoustic model discovery
C. Lee and J. R. Glass · 2012
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Generative adversarial networks
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Earlier work this paper cites.
Using different acoustic, lexical and language modeling units for asr of an under-resourced language - amharic
M. Tachbelie, S. T. Abate, and L. Besacier · 2014
Earlier work this paper cites.
Speech technologies for african languages: example of a multilingual calculator for education
L. Besacier, E. Gauthier, M. Mangeot, P. Bretier, P. Bagshaw, O. Rosec, T. Moudenc, F. Pellegrino, S. Voisin, E. Marsico, and P. Nocera · 2015
Earlier work this paper cites.
Attention-based models for speech recognition
J. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Unsupervised lexicon discovery from acoustic input
C. Lee, T. J. O’Donnell, and J. R. Glass · 2015
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur · 2015
Cited alongside, same era.
Unsupervised word discovery from speech using automatic segmentation into syllable-like units
O. Rasanen, G. Doyle, and M. C. Frank · 2015
Cited alongside, same era.
Improving neural machine translation models with monolingual data
R. Sennrich, B. Haddow, and A. Birch · 2015
Cited alongside, same era.
Deep speech 2: End-to-end speech recognition in english and mandarin
D. Amodei, S. Ananthanarayanan, R. Anubhai, J. Bai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, Q. Cheng, G. Chen, et al · 2016
Cited alongside, same era.
Domain-adversarial training of neural networks
Y. Ganin, E. Ustinova, H. Ajakan, P. Germain, H. Larochelle, F. Laviolette, M. Marchand, and V. Lempitsky · 2016
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Completely unsupervised speech recognition by a generative adversarial network harmonized with iteratively refined hidden markov models
K.-Y. Chen, C.-P. Tsai, D.-R. Liu, H.-Y. Lee, and L. shan Lee · 2019
Later among the works it cites.
Cross-lingual language model pretraining
A. Conneau and G. Lample · 2019
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Later among the works it cites.
Towards visually grounded sub-word speech unit discovery
D. Harwath and J. Glass · 2019
Later among the works it cites.
Cycle-consistency training for end-to-end speech recognition
T. Hori, R. Astudillo, T. Hayashi, Y. Zhang, S. Watanabe, and J. Le Roux · 2019
Later among the works it cites.
What does bert learn about the structure of language?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Jang, S. Gu, and B. Poole · 2016
Cited alongside, same era.
Ethnologue: Languages of the world, nineteenth edition
M. P. Lewis, G. F. Simon, and C. D. Fennig · 2016
Cited alongside, same era.
Variational inference for acoustic unit discovery
L. Ondel, L. Burget, and J. Cernocký · 2016
Cited alongside, same era.
Dual learning for machine translation
Y. Xia, D. He, T. Qin, L. Wang, N. Yu, T. Liu, and W. Ma · 2016
Cited alongside, same era.
Wasserstein gan
M. Arjovsky, S. Chintala, and L. Bottou · 2017
Cited alongside, same era.
Learning bilingual word embeddings with (almost) no bilingual data
M. Artetxe, G. Labaka, and E. Agirre · 2017
Cited alongside, same era.
Towards better decoding and language model integration in sequence to sequence models
J. Chorowski and N. Jaitly · 2017
Cited alongside, same era.
G. Jawahar, B. Sagot, and D. Seddah · 2019
Later among the works it cites.
Improving transformer-based speech recognition using unsupervised pre-training
D. Jiang, X. Lei, W. Li, N. Luo, Y. Hu, W. Zou, and X. Li · 2019
Later among the works it cites.
Billion-scale similarity search with gpus
J. Johnson, M. Douze, and H. Jégou · 2019
Later among the works it cites.
Towards unsupervised speech recognition and synthesis with quantized speech representation learning
A. H. Liu, T. Tu, H. yi Lee, and L. shan Lee · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli · 2019
Later among the works it cites.
Specaugment: A simple data augmentation method for automatic speech recognition
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le · 2019
Later among the works it cites.
Wav2letter++: A fast open-source speech recognition system
V. Pratap, A. Hannun, Q. Xu, J. Cai, J. Kahn, G. Synnaeve, V. Liptchinsky, and R. Collobert · 2019
Later among the works it cites.
The pytorch-kaldi speech recognition toolkit
M. Ravanelli, T. Parcollet, and Y. Bengio · 2019
Later among the works it cites.
wav2vec: Unsupervised pre-training for speech recognition
S. Schneider, A. Baevski, R. Collobert, and M. Auli · 2019
Later among the works it cites.
Bert rediscovers the classical nlp pipeline
I. Tenney, D. Das, and E. Pavlick · 2019
Later among the works it cites.
Unsupervised speech recognition via segmental empirical output distribution matching
C.-K. Yeh, J. Chen, C. Yu, and D. Yu · 2019
Later among the works it cites.
Common voice: A massively-multilingual speech corpus
R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber · 2020
Later among the works it cites.
Unsupervised cross-lingual representation learning for speech recognition
A. Conneau, A. Baevski, R. Collobert, A. Mohamed, and M. Auli · 2020
Later among the works it cites.
Conformer: Convolution-augmented transformer for speech recognition
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang · 2020
Later among the works it cites.
Contextnet: Improving convolutional neural networks for automatic speech recognition with global context
W. Han, Z. Zhang, Y. Zhang, J. Yu, C.-C. Chiu, J. Qin, A. Gulati, R. Pang, and Y. Wu · 2020
Later among the works it cites.
Differentiable weighted finite-state transducers
A. Hannun, V. Pratap, J. Kahn, and W.-N. Hsu · 2020
Later among the works it cites.
Learning hierarchical discrete linguistic units from visually-grounded speech
D. Harwath, W.-N. Hsu, and J. Glass · 2020
Later among the works it cites.
Semi-supervised speech recognition via local prior matching
W.-N. Hsu, A. Lee, G. Synnaeve, and A. Hannun · 2020
Later among the works it cites.
Learning robust and multilingual speech representations
K. Kawakami, L. Wang, C. Dyer, P. Blunsom, and A. van den Oord · 2020
Later among the works it cites.
Self-supervised contrastive learning for unsupervised phoneme segmentation
F. Kreuk, J. Keshet, and Y. Adi · 2020
Later among the works it cites.
Improved noisy student training for automatic speech recognition
D. S. Park, Y. Zhang, Y. Jia, W. Han, C.-C. Chiu, B. Li, Y. Wu, and Q. V. Le · 2020
Later among the works it cites.
Mls: A large-scale multilingual dataset for speech research
V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert · 2020
Later among the works it cites.
Unsupervised pretraining transfers well across languages
M. Rivière, A. Joulin, P.-E. Mazaré, and E. Dupoux · 2020
Later among the works it cites.
End-to-end ASR: from Supervised to Semi-Supervised Learning with Modern Architectures
G. Synnaeve, Q. Xu, J. Kahn, T. Likhomanenko, E. Grave, V. Pratap, A. Sriram, V. Liptchinsky, and R. Collobert · 2020
Later among the works it cites.
rvad: An unsupervised segment-based robust voice activity detection method
Z. Tan, A. K. Sarkar, and N. Dehak · 2020
Later among the works it cites.
Vector-quantized neural networks for acoustic unit discovery in the zerospeech 2020 challenge
B. van Niekerk, L. Nortje, and H. Kamper · 2020
Later among the works it cites.
Exploring wav2vec 2.0 on speaker verification and language identification
Z. Fan, M. Li, S. Zhou, and B. Xu · 2021
Closest in time.
Google cloud: Speech-to-text
Google · 2021
Closest in time.
slimipl: Language-model-free iterative pseudo-labeling
T. Likhomanenko, Q. Xu, J. Kahn, G. Synnaeve, and R. Collobert · 2021
Closest in time.
Emotion recognition from speech using wav2vec 2.0 embeddings
L. Pepino, P. Riera, and L. Ferrer · 2021
Closest in time.
Large-scale self- and semi-supervised learning for speech translation
C. Wang, A. Wu, J. Pino, A. Baevski, M. Auli, and A. Conneau · 2021
Closest in time.