Fetching the paper…
Reading the bibliography…
We introduce a new collection of spoken English audio suitable for training speech recognition systems under limited or no supervision.
“Active learning for automatic speech recognition,”
D. Hakkani-Tür, G. Riccardi, and A. Gorin, · 2002
Earlier work this paper cites.
“Combining active and semi-supervised learning for spoken language understanding,”
G. Tur, D. Hakkani-Tür, and R.E. Schapire, · 2005
Earlier work this paper cites.
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Kenlm: Faster and smaller language model queries,”
Kenneth Heafield, · 2011
Earlier work this paper cites.
“The language-independent bottleneck features,”
K. Veselỳ, M. Karafiát, F. Grézl, M. Janda, and E. Egorova, · 2012
Earlier work this paper cites.
“Cross-language knowledge transfer using multilingual deep neural network with shared hidden layers,”
J.-T. Huang, J. Li, D. Yu, L. Deng, and Y. Gong, · 2013
Earlier work this paper cites.
“Evaluating speech features with the minimal-pair abx task: Analysis of the classical mfc/plp pipeline,”
T. Schatz, V. Peddinti, F. Bach, A. Jansen, H. Hermansky, and E. Dupoux, · 2013
Earlier work this paper cites.
“Iarpa babel program,” 2014
M. Harper, · 2014
Earlier work this paper cites.
“Distant supervision for representation learning in speech and handwriting,” 2014
J. Chorowski, · 2014
Earlier work this paper cites.
“Librispeech: an asr
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“The zero resource speech challenge 2015: Proposed approaches and results,”
M. Versteegh, X. Anguera, A. Jansen, and E. Dupoux, · 2016
Cited alongside, same era.
ABX-discriminability measures and applications
T. Schatz, · 2016
Cited alongside, same era.
“The zero resource speech challenge 2017,”
E. Dunbar, X.-N. Cao, J. Benjumea, J. Karadayi, M. Bernard, L. Besacier, X. Anguera, and E. Dupoux, · 2017
Cited alongside, same era.
“An empirical evaluation of zero resource acoustic unit discovery,”
C. Liu, J. Yang, M. Sun, S. Kesiraju, A. Rott, L. Ondel, P. Ghahremani, N. Dehak, L. Burget, and S. Khudanpur, · 2017
Cited alongside, same era.
“Letter based speech recognition with gated convnets,”
V. Liptchinsky, G. Synnaeve, and R. Collobert, · 2017
Cited alongside, same era.
“Feature optimized DPGMM clustering for unsupervised subword modeling: A contribution to zerospeech 2017,”
“Speech2vec: A sequence-to-sequence framework for learning word embeddings from speech,”
Y.-A. Chung and J. Glass, · 2018
Later among the works it cites.
“Self-training for end-to-end speech recognition,”
J. Kahn, A. Lee, and A. Hannun, · 2019
Closest in time.
“wav2vec: Unsupervised pre-training for speech recognition,”
S. Schneider, A. Baevski, R. Collobert, and M. Auli, · 2019
Closest in time.
“The zero resource speech challenge 2019: TTS without T,”
E. Dunbar, R. Algayres, J. Karadayi, M. Bernard, J. Benjumea, X.-N. Cao, L. Miskic, C. Dugrain, L. Ondel, A. Black, L. Besacier, S. Sakriani, and E. Dupoux, · 2019
Closest in time.
“CMU wilderness multilingual speech dataset,”
A.W. Black, · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Heck, S. Sakti, and S. Nakamura, · 2017
Cited alongside, same era.
“Almost-unsupervised speech recognition with close-to-zero resource based on phonetic structures learned from very small unpaired speech and text data,”
Y.-C. Chen, C.-H. Shen, S.-F. Huang, H.-y. Lee, and L.-s. Lee, · 2018
Cited alongside, same era.
“Unsupervised cross-modal alignment of speech and text embedding spaces,”
Y.-A. Chung, W.-H. Weng, S. Tong, and J. Glass, · 2018
Cited alongside, same era.
“Representation learning with contrastive predictive coding,”
A. van den Oord, Y. Li, and O. Vinyals, · 2018
Cited alongside, same era.
“Wav2letter++: A fast open-source speech recognition system,”
V. Pratap, A. Hannun, Q. Xu, J. Cai, J. Kahn, G. Synnaeve, V. Liptchinsky, and R. Collobert, · 2019
Closest in time.
“Rwth asr systems for librispeech: Hybrid vs attention-w/o data augmentation,”
C. Lüscher, E. Beck, K. Irie, M. Kitza, W. Michel, A. Zeyer, R. Schlüter, and H. Ney, · 2019
Closest in time.
“Sequence-to-sequence speech recognition with time-depth separable convolutions,”
A. Hannun, A. Lee, Q. Xu, and R. Collobert, · 2019
Closest in time.
“vq-wav2vec: Self-supervised learning of discrete speech representations,”
Anonymous, · 2020
Closest in time.