Fetching the paper…
Reading the bibliography…
Self-supervised approaches for speech representation learning are challenged by three unique problems: (1) there are multiple sound units in each input utterance, (2) there is no lexicon of input sound units during the pre-training phase, and (3) sound units have variable lengths with no explicit segmentation.
S. Lloyd, “Least squares quantization in pcm,” IEEE transactions on information theory , vol. 28, no. 2, pp. 129–137, 1982
1982
Earlier work this paper cites.
S. Young, “Large vocabulary continuous speech recognition: A review,” IEEE Signal Processing Magazine , vol. 13, no. 5, pp. 45–57, 1996
1996
Earlier work this paper cites.
G. Zavaliagkos and T. Colthurst, “Utilizing untranscribed training data to improve performance,” in DARPA Broadcast News Transcription and Understanding Workshop , 1998
1998
Earlier work this paper cites.
R. M. Gray and D. L. Neuhoff, “Quantization,” IEEE transactions on information theory , vol. 44, no. 6, pp. 2325–2383, 1998
1998
Earlier work this paper cites.
D. Povey, “Discriminative training for large vocabulary speech recognition,” Ph.D. dissertation, University of Cambridge, 2005
2005
Earlier work this paper cites.
J. Ma, S. Matsoukas, O. Kimball, and R. Schwartz, “Unsupervised training on large amounts of broadcast news data,” in ICASSP , 2006
2006
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in ICML , 2006
2006
Earlier work this paper cites.
D. Arthur and S. Vassilvitskii, “k-means++: The advantages of careful seeding,” Stanford, Tech. Rep., 2006
2006
Earlier work this paper cites.
F. Pedregosa et al. , “Scikit-learn: Machine learning in python,” the Journal of machine Learning research , 2011
2011
Earlier work this paper cites.
C.-y. Lee and J. Glass, “A nonparametric bayesian approach to acoustic model discovery,” in ACL , 2012
2012
Earlier work this paper cites.
O. Abdel-Hamid, A.-r. Mohamed, H. Jiang, and G. Penn, “Applying convolutional neural networks concepts to hybrid nn-hmm model for speech recognition,” in 2012 IEEE international conference on Acoustics, speech and signal processing (ICASSP) . IEEE, 2012, pp. 4277–4280
2012
Earlier work this paper cites.
H. A. Bourlard and N. Morgan, Connectionist speech recognition: a hybrid approach . Springer Science & Business Media, 2012, vol. 247
2012
Earlier work this paper cites.
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in ICASSP , 2015
2015
Earlier work this paper cites.
L. Ondel, L. Burget, and J. Černockỳ, “Variational inference for acoustic unit discovery,” Procedia Computer Science , vol. 81, pp. 80–86, 2016
2016
Earlier work this paper cites.
J. Ebbers, J. Heymann, L. Drude, T. Glarner, R. Haeb-Umbach, and B. Raj, “Hidden markov model variational autoencoder for acoustic unit discovery.” in INTERSPEECH , 2017
2017
Earlier work this paper cites.
W.-N. Hsu, Y. Zhang, and J. Glass, “Learning latent representations for speech generation and transformation,” in INTERSPEECH , 2017
2017
Earlier work this paper cites.
——, “Unsupervised learning of disentangled and interpretable representations from sequential data,” in NeurIPS , 2017
2017
Earlier work this paper cites.
A. van den Oord, O. Vinyals et al. , “Neural discrete representation learning,” in NeurIPS , 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer, “Deep contextualized word representations,” in NAACL , 2018
2018
Earlier work this paper cites.
M. Caron, P. Bojanowski, A. Joulin, and M. Douze, “Deep clustering for unsupervised learning of visual features,” in ECCV , 2018
2018
Cited alongside, same era.
T. Glarner, P. Hanebrink, J. Ebbers, and R. Haeb-Umbach, “Full bayesian hidden markov model variational autoencoder for acoustic unit discovery.” in INTERSPEECH , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Cited alongside, same era.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in CVPR , 2020
2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
J. Chorowski, R. J. Weiss, S. Bengio, and A. van den Oord, “Unsupervised speech representation learning using wavenet autoencoders,” IEEE/ACM transactions on audio, speech, and language processing , vol. 27, no. 12, pp. 2041–2053, 2019
2019
Cited alongside, same era.
S. Khurana, S. R. Joty, A. Ali, and J. Glass, “A factorial deep markov model for unsupervised disentangled representation learning from speech,” in ICASSP , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
S. Pascual, M. Ravanelli, J. Serrà, A. Bonafonte, and Y. Bengio, “Learning problem-agnostic speech representations from multiple self-supervised tasks,” in INTERSPEECH , 2019
2019
Cited alongside, same era.
Later among the works it cites.
J. Kahn et al. , “Libri-light: A benchmark for asr with limited or no supervision,” in ICASSP , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
M. Joshi, D. Chen, Y. Liu, D. S. Weld, L. Zettlemoyer, and O. Levy, “Spanbert: Improving pre-training by representing and predicting spans,” Transactions of the Association for Computational Linguistics , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
Y.-A. Chung and J. Glass, “Generative pre-training for speech with autoregressive predictive coding,” in ICASSP , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
S. Ling, Y. Liu, J. Salazar, and K. Kirchhoff, “Deep contextualized acoustic representations for semi-supervised speech recognition,” in ICASSP , 2020
2020
Later among the works it cites.
W. Wang, Q. Tang, and K. Livescu, “Unsupervised pre-training of bidirectional speech encoders via masked reconstruction,” in ICASSP , 2020
2020
Later among the works it cites.
A. T. Liu, S.-w. Yang, P.-H. Chi, P.-c. Hsu, and H.-y. Lee, “Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders,” in ICASSP , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.