Fetching the paper…
Reading the bibliography…
Self-supervised learning of speech representations has been a very active research area but most work is focused on a single domain such as read audio books for which there exist large quantities of labeled and unlabeled data.
Switchboard-1 Release 2 LDC97S62
John Godfrey and Edward Holliman · 1993
Earlier work this paper cites.
CSR-I (WSJ0) complete LDC93S6A
John Garofalo, David Graff, Doug Paul, and David Pallett · 1993
Earlier work this paper cites.
CSR-II (WSJ1) Complete LDC94S13A
Linguistic Data Consortium, NIST Multimodal Information Group · 1994
Earlier work this paper cites.
Robust speech recognition using the modulation spectrogram
Brian ED Kingsbury, Nelson Morgan, and Steven Greenberg · 1998
Earlier work this paper cites.
Fisher English training speech parts 1 and 2 LDC200{4,5}S13
Cieri, Christopher, et al. · 2005
Earlier work this paper cites.
Fisher English training speech parts 1 and 2 transcripts LDC200{4,5}T19
Cieri, Christopher, et al. · 2005
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber · 2006
Earlier work this paper cites.
2003 NIST Rich Transcription Evaluation Data LDC2007S10
Fiscus, Jonathan G., et al · 2007
Earlier work this paper cites.
2000 HUB5 English Evaluation Speech LDC2002S09
Linguistic Data Consortium · 2007
Earlier work this paper cites.
Product quantization for nearest neighbor search
H. Jegou, M. Douze, and C. Schmid · 2011
Earlier work this paper cites.
Kenlm: Faster and smaller language model queries
Kenneth Heafield · 2011
Earlier work this paper cites.
Features based on auditory physiology and perception
Richard M Stern, Nelson Morgan, T Virtanen, B Raj, and R Singh · 2012
Earlier work this paper cites.
An investigation of deep neural networks for noise robust speech recognition
Michael L Seltzer, Dong Yu, and Yongqiang Wang · 2013
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur · 2015
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Earlier work this paper cites.
Layer normalization
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Cited alongside, same era.
Unsupervised learning of disentangled and interpretable representations from sequential data
Wei-Ning Hsu, Yu Zhang, and James Glass · 2017
Cited alongside, same era.
Generation of large-scale simulated utterances in virtual rooms to train deep-neural networks for far-field speech recognition in google home
Chanwoo Kim, Ananya Misra, Kean Chin, Thad Hughes, Arun Narayanan, Tara Sainath, and Michiel Bacchiani · 2017
Cited alongside, same era.
An unsupervised deep domain adaptation approach for robust speech recognition
Sining Sun, Binbin Zhang, Lei Xie, and Yanning Zhang · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, and et al · 2017
Cited alongside, same era.
Wav2letter++: A fast open-source speech recognition system
Vineel Pratap, Awni Hannun, Qiantong Xu, Jeff Cai, Jacob Kahn, Gabriel Synnaeve, Vitaliy Liptchinsky, and Ronan Collobert · 2019
Later among the works it cites.
Learning hierarchical discrete linguistic units from visually-grounded speech
David Harwath, Wei-Ning Hsu, and James Glass · 2020
Later among the works it cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
A. Baevski, H. Zhou, A. Mohamed, and M. Auli · 2020
Later among the works it cites.
Rethinking evaluation in asr: Are our models robust enough?
Tatiana Likhomanenko, Qiantong Xu, Vineel Pratap, Paden Tomasello, Jacob Kahn, Gilad Avidov, Ronan Collobert, and Gabriel Synnaeve · 2020
Later among the works it cites.
Unsupervised domain adaptation for speech recognition via uncertainty driven self-training
Sameer Khurana, Niko Moritz, Takaaki Hori, and Jonathan Le Roux · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. v. d. Oord, Y. Li, and O. Vinyals · 2018
Cited alongside, same era.
Extracting domain invariant features by unsupervised learning for robust automatic speech recognition
Wei-Ning Hsu and James Glass · 2018
Cited alongside, same era.
Hao Tang, Wei-Ning Hsu, François Grondin, and James Glass · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Ted-lium 3: twice as much data and corpus repartition for experiments on speaker adaptation
François Hernandez, Vincent Nguyen, Sahar Ghannay, Natalia Tomashenko, and Yannick Estève · 2018
Cited alongside, same era.
wav2vec: Unsupervised pre-training for speech recognition
S. Schneider, A. Baevski, R. Collobert, and M. Auli · 2019
Cited alongside, same era.
An unsupervised autoregressive model for speech representation learning
Y.-A. Chung, W.-N. Hsu, H. Tang, and J. R. Glass · 2019
Cited alongside, same era.
Learning robust and multilingual speech representations
K. Kawakami, L. Wang, C. Dyer, P. Blunsom, and A. v. d. Oord · 2020
Later among the works it cites.
Unsupervised pretraining transfers well across languages
M. Rivière, A. Joulin, P.-E. Mazaré, and E. Dupoux · 2020
Later among the works it cites.
Unsupervised cross-lingual representation learning for speech recognition
Alexis Conneau, Alexei Baevski, Ronan Collobert, Abdelrahman Mohamed, and Michael Auli · 2020
Later among the works it cites.
vq-wav2vec: Self-supervised learning of discrete speech representations
A. Baevski, S. Schneider, and M. Auli · 2020
Later among the works it cites.
Libri-light: A benchmark for asr with limited or no supervision
J. Kahn and et al · 2020
Later among the works it cites.
Reducing transformer depth on demand with structured dropout
A. Fan, E. Grave, and A. Joulin · 2020
Later among the works it cites.
Conformer: Convolution-augmented transformer for speech recognition
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, and et al · 2020
Later among the works it cites.
An investigation of phone-based subword units for end-to-end speech recognition
Weiran Wang, Yingbo Zhou, Caiming Xiong, and Richard Socher · 2020
Later among the works it cites.
The rwth asr system for ted-lium release 2: Improving hybrid hmm with specaugment
Wei Zhou, Wilfried Michel, Kazuki Irie, Markus Kitza, Ralf Schlüter, and Hermann Ney · 2020
Later among the works it cites.
Changhan Wang, Morgane Rivière, Ann Lee, Anne Wu, Chaitanya Talnikar, Daniel Haziza, Mary Williamson, Juan Pino, and Emmanuel Dupoux · 2021
Closest in time.