Fetching the paper…
Reading the bibliography…
Unsupervised representation learning for speech processing has matured greatly in the last few years.
Improving transformer-based speech recognition using unsupervised pre-training
Jiang, D.; Lei, X.; Li, W.; Luo, N.; Hu, Y.; Zou, W.; and Li, X. 2019 · 1910
Earlier work this paper cites.
Effectiveness of self-supervised pre-training for speech recognition
Baevski, A.; Auli, M.; and Mohamed, A. 2019 · 1911
Earlier work this paper cites.
A generalization of isolated word recognition using vector quantization
Burton, D. K.; Shore, J. E.; and Buck, J. T. 1983 · 1983
Earlier work this paper cites.
A vector Quantization approach to Speaker Recognition
Soong, F.; Rosenberg, A.; and Juang, L. R. B. 1985 · 1985
Earlier work this paper cites.
Chapter 6: Information Processing in Dynamical Systems: Foundations of Harmony Theory
Smolensky, P. 1986 · 1986
Earlier work this paper cites.
Nonlinear principal component analysis using autoassociative neural networks
Kramer, M. A. 1991 · 1991
Earlier work this paper cites.
TIMIT Acoustic-Phonetic Continuous Speech Corpus LDC93S1
Garofolo, J. S. 1993 · 1993
Earlier work this paper cites.
An Introduction to Variational Methods for Graphical Models
Jordan, M. I.; Ghahramani, Z.; Jaakkola, T. S.; and Saul, L. K. 1999 · 1999
Earlier work this paper cites.
Slow feature analysis: Unsupervised learning of invariances
Wiskott, L.; and Sejnowski, T. J. 2002 · 2002
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Graves, A.; Fernández, S.; Gomez, F.; and Schmidhuber, J. 2006 · 2006
Earlier work this paper cites.
A Fast Learning Algorithm for Deep Belief Nets
Hinton, G. E.; Osindero, S.; and Teh, Y. W. 2006 · 2006
Earlier work this paper cites.
An overview of deep semi-supervised learning
Ouali, Y.; Hudelot, C.; and Tami, M. 2020 · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Unsupervised feature learning for audio classification using convolutional deep belief networks
Lee, H.; Largman, Y.; Pham, P.; and Ng, A. Y. 2009 · 2009
Earlier work this paper cites.
SPLAT: Speech-Language Joint Pre-Training for Spoken Language Understanding
Chung, Y.-A.; Zhu, C.; and Zeng, M. 2020 · 2010
Earlier work this paper cites.
Binary coding of speech spectrograms using a deep auto-encoder
Deng, L.; Seltzer, M. L.; Yu, D.; Acero, A.; Mohamed, A.; and Hinton, G. E. 2010 · 2010
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M.; and Hyvärinen, A. 2010 · 2010
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, T.; Karafiát, M.; Burget, L.; Cernockỳ, J.; and Khudanpur, S. 2010 · 2010
Earlier work this paper cites.
Latent variable models and factor analysis: a unified approach
Bartholomew, D. J.; Knott, M.; and Moustaki, I. 2011 · 2011
Earlier work this paper cites.
Bounding the Bias of Contrastive Divergence Learning
Fischer, A.; and Igel, C. 2011 · 2011
Earlier work this paper cites.
A Practical Guide to Training Restricted Boltzmann Machines
Hinton, G. E. 2012 · 2012
Earlier work this paper cites.
A Nonparametric Bayesian Approach to Acoustic Model Discovery
Lee, C.-y.; and Glass, J. 2012 · 2012
Earlier work this paper cites.
Decoar 2.0: Deep contextualized acoustic representations with vector quantization
Ling, S.; and Liu, Y. 2020 · 2012
Earlier work this paper cites.
Representation Learning: A Review and New Perspectives
Bengio, Y.; Courville, A. C.; and Vincent, P. 2013 · 2013
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y.; Léonard, N.; and Courville, A. 2013 · 2013
Earlier work this paper cites.
Weak top-down constraints for unsupervised acoustic model training
Jansen, A.; Thomas, S.; and Hermansky, H. 2013 · 2013
Earlier work this paper cites.
Fixed-dimensional acoustic embeddings of variable-length segments in low-resource settings
Levin, K.; Henry, K.; Jansen, A.; and Livescu, K. 2013 · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T.; Sutskever, I.; Chen, K.; Corrado, G. S.; and Dean, J. 2013 · 2013
Earlier work this paper cites.
Evaluating speech features with the minimal-pair ABX task: Analysis of the classical MFC/PLP pipeline
Schatz, T.; Peddinti, V.; Bach, F.; Jansen, A.; Hermansky, H.; and Dupoux, E. 2013 · 2013
Earlier work this paper cites.
An auto-encoder based approach to unsupervised learning of subword units
Badino, L.; Canevari, C.; Fadiga, L.; and Metta, G. 2014 · 2014
Earlier work this paper cites.
Training Restricted Boltzmann Machines: An Introduction
Fischer, A.; and Igel, C. 2014 · 2014
Earlier work this paper cites.
Auto-Encoding Variational Bayes
Kingma, D. P.; and Welling, M. 2014 · 2014
Earlier work this paper cites.
Stochastic Backpropagation and Approximate Inference in Deep Generative Models
Rezende, D. J.; Mohamed, S.; and Wierstra, D. 2014 · 2014
Earlier work this paper cites.
Evaluating speech features with the Minimal-Pair ABX task (II): Resistance to noise
Schatz, T.; Peddinti, V.; Cao, X.-N.; Bach, F.; Hermansky, H.; and Dupoux, E. 2014 · 2014
Earlier work this paper cites.
A Recurrent Latent Variable Model for Sequential Data
Chung, J.; Kastner, K.; Dinh, L.; Goel, K.; Courville, A. C.; and Bengio, Y. 2015 · 2015
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
Doersch, C.; Gupta, A.; and Efros, A. A. 2015 · 2015
Earlier work this paper cites.
Unsupervised neural network based feature extraction using weak top-down constraints
Kamper, H.; Elsner, M.; Jansen, A.; and Goldwater, S. 2015 · 2015
Earlier work this paper cites.
A comparison of neural network methods for unsupervised representation learning on the Zero Resource Speech Challenge
Renshaw, D.; Kamper, H.; Jansen, A.; and Goldwater, S. 2015 · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K.; and Zisserman, A. 2015 · 2015
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C.; Liu, W.; Jia, Y.; Sermanet, P.; Reed, S.; Anguelov, D.; Erhan, D.; Vanhoucke, V.; and Rabinovich, A. 2015 · 2015
Earlier work this paper cites.
The Zero Resource Speech Challenge 2015
Versteegh, M.; Thiolliere, R.; Schatz, T.; Cao, X. N.; Anguera, X.; Jansen, A.; and Dupoux, E. 2015 · 2015
Earlier work this paper cites.
Generating Sentences from a Continuous Space
Bowman, S. R.; Vilnis, L.; Vinyals, O.; Dai, A. M.; Józefowicz, R.; and Bengio, S. 2016 · 2016
Cited alongside, same era.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
Chan, W.; Jaitly, N.; Le, Q.; and Vinyals, O. 2016 · 2016
Cited alongside, same era.
Audio word2vec: Unsupervised learning of audio segment representations using sequence-to-sequence autoencoder
Chung, Y.-A.; Wu, C.-C.; Shen, C.-H.; Lee, H.-Y.; and Lee, L.-S. 2016 · 2016
Cited alongside, same era.
Sequential Neural Models with Stochastic Layers
Fraccaro, M.; Sønderby, S. K.; Paquet, U.; and Winther, O. 2016 · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Cited alongside, same era.
Improved Variational Inference with Inverse Autoregressive Flow
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
Baevski, A.; Zhou, Y.; Mohamed, A.; and Auli, M. 2020 · 2020
Later among the works it cites.
Vector-Quantized Autoregressive Predictive Coding
Chung, Y.-A.; Tang, H.; and Glass, J. 2020 · 2020
Later among the works it cites.
ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators
Clark, K.; Luong, M.; Le, Q. V.; and Manning, C. D. 2020 · 2020
Later among the works it cites.
Unsupervised cross-lingual representation learning for speech recognition
Conneau, A.; Baevski, A.; Collobert, R.; Mohamed, A.; and Auli, M. 2020 · 2020
Later among the works it cites.
The Zero Resource Speech Challenge 2020: Discovering discrete subword and word units
Dunbar, E.; Karadayi, J.; Bernard, M.; Cao, X.-N.; Algayres, R.; Ondel, L.; Besacier, L.; Sakti, S.; and Dupoux, E. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kingma, D. P.; Salimans, T.; Jozefowicz, R.; Chen, X.; Sutskever, I.; and Welling, M. 2016 · 2016
Cited alongside, same era.
Variational Inference for Acoustic Unit Discovery
Ondel, L.; Burget, L.; and Cernocký, J. 2016 · 2016
Cited alongside, same era.
Context encoders: Feature learning by inpainting
Pathak, D.; Krahenbuhl, P.; Donahue, J.; Darrell, T.; and Efros, A. A. 2016 · 2016
Cited alongside, same era.
Ladder Variational Autoencoders
Sønderby, C. K.; Raiko, T.; Maaløe, L.; Sønderby, S. K.; and Winther, O. 2016 · 2016
Cited alongside, same era.
WaveNet: A Generative Model for Raw Audio
van den Oord, A.; Dieleman, S.; Zen, H.; Simonyan, K.; Vinyals, O.; Graves, A.; Kalchbrenner, N.; Senior, A.; and Kavukcuoglu, K. 2016 · 2016
Cited alongside, same era.
Towards speech-to-text translation without speech recognition
Bansal, S.; Kamper, H.; Lopez, A.; and Goldwater, S. 2017 · 2017
Cited alongside, same era.
Learning word embeddings from speech
Chung, Y.-A.; and Glass, J. 2017 · 2017
Cited alongside, same era.
Gulati, A.; Qin, J.; Chiu, C.-C.; Parmar, N.; Zhang, Y.; Yu, J.; Han, W.; Wang, S.; Zhang, Z.; Wu, Y.; et al. 2020 · 2020
Later among the works it cites.
Is the Discrete VAE’s Power Stuck in its Prior?
Jones, H. T.; and Moore, J. 2020 · 2020
Later among the works it cites.
Learning Robust and Multilingual Speech Representations
Kawakami, K.; Wang, L.; Dyer, C.; Blunsom, P.; and van den Oord, A. 2020 · 2020
Later among the works it cites.
A Convolutional Deep Markov Model for Unsupervised Speech Representation Learning
Khurana, S.; Laurent, A.; Hsu, W.-N.; Chorowski, J.; Lancucki, A.; Marxer, R.; and Glass, J. 2020 · 2020
Later among the works it cites.
Self-Supervised Contrastive Learning for Unsupervised Phoneme Segmentation
Kreuk, F.; Keshet, J.; and Adi, Y. 2020 · 2020
Later among the works it cites.
Deep contextualized acoustic representations for semi-supervised speech recognition
Ling, S.; Liu, Y.; Salazar, J.; and Kirchhoff, K. 2020 · 2020
Later among the works it cites.
Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders
Liu, A. T.; Yang, S.-w.; Chi, P.-H.; Hsu, P.-c.; and Lee, H.-y. 2020 · 2020
Later among the works it cites.
Monte Carlo Gradient Estimation in Machine Learning
Mohamed, S.; Rosca, M.; Figurnov, M.; and Mnih, A. 2020 · 2020
Later among the works it cites.
Multi-task self-supervised learning for Robust Speech Recognition
Ravanelli, M.; Zhong, J.; Pascual, S.; Swietojanski, P.; Monteiro, J.; Trmal, J.; and Bengio, Y. 2020 · 2020
Later among the works it cites.
Unsupervised pretraining transfers well across languages
Riviere, M.; Joulin, A.; Mazaré, P.-E.; and Dupoux, E. 2020 · 2020
Later among the works it cites.
Pre-training audio representations with self-supervision
Tagliasacchi, M.; Gfeller, B.; de Chaumont Quitry, F.; and Roblek, D. 2020 · 2020
Later among the works it cites.
Self-supervised Learning from a Multi-view Perspective
Tsai, Y.-H. H.; Wu, Y.; Salakhutdinov, R.; and Morency, L.-P. 2020 · 2020
Later among the works it cites.
Vector-Quantized Neural Networks for Acoustic Unit Discovery in the ZeroSpeech 2020 Challenge
van Niekerk, B.; Nortje, L.; and Kamper, H. 2020 · 2020
Later among the works it cites.
Unsupervised pre-training of bidirectional speech encoders via masked reconstruction
Wang, W.; Tang, Q.; and Livescu, K. 2020 · 2020
Later among the works it cites.
Unsupervised Speech Recognition
Baevski, A.; Hsu, W.-N.; Conneau, A.; and Auli, M. 2021 · 2021
Later among the works it cites.
WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
Chen, S.; Wang, C.; Chen, Z.; Wu, Y.; Liu, S.; Chen, Z.; Li, J.; Kanda, N.; Yoshioka, T.; Xiao, X.; et al. 2021 · 2021
Later among the works it cites.
Audio albert: A lite bert for self-supervised learning of audio representation
Chi, P.-H.; Chung, P.-H.; Wu, T.-H.; Hsieh, C.-C.; Chen, Y.-H.; Li, S.-W.; and Lee, H.-y. 2021 · 2021
Later among the works it cites.
Cuervo, S.; Grabias, M.; Chorowski, J.; Ciesielski, G.; Łańcucki, A.; Rychlikowski, P.; and Marxer, R. 2021 · 2021
Later among the works it cites.
Variable-rate discrete representation learning
Dieleman, S.; Nash, C.; Engel, J.; and Simonyan, K. 2021 · 2021
Later among the works it cites.
The Zero Resource Speech Challenge 2021: Spoken language modelling
Dunbar, E.; Bernard, M.; Hamilakis, N.; Nguyen, T.; de Seyssel, M.; Rozé, P.; Rivière, M.; Kharitonov, E.; and Dupoux, E. 2021 · 2021
Later among the works it cites.
HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
Hsu, W.-N.; Bolte, B.; Tsai, Y.-H. H.; Lakhotia, K.; Salakhutdinov, R.; and Mohamed, A. 2021 · 2021
Later among the works it cites.
not-MIWAE: Deep generative modelling with missing not at random data
Ipsen, N. B.; Mattei, P.-A.; and Frellsen, J. 2021 · 2021
Later among the works it cites.
Acoustic word embeddings for zero-resource languages using self-supervised contrastive learning and multilingual adaptation
Jacobs, C.; Matusevych, Y.; and Kamper, H. 2021 · 2021
Later among the works it cites.
A Further Study of Unsupervised Pretraining for Transformer Based Speech Recognition
Jiang, D.; Li, W.; Zhang, R.; Cao, M.; Luo, N.; Han, Y.; Zou, W.; Han, K.; and Li, X. 2021 · 2021
Later among the works it cites.
Towards unsupervised phone and word segmentation using self-supervised vector-quantized neural networks
Kamper, H.; and van Niekerk, B. 2021 · 2021
Later among the works it cites.
Magic dust for cross-lingual adaptation of monolingual wav2vec-2.0
Khurana, S.; Laurent, A.; and Glass, J. 2021 · 2021
Later among the works it cites.
Semi-supervised spoken language understanding via self-supervised speech and language model pretraining
Lai, C.-I.; Chuang, Y.-S.; Lee, H.-Y.; Li, S.-W.; and Glass, J. 2021 · 2021
Later among the works it cites.
Non-autoregressive predictive coding for learning speech representations from local dependencies
Liu, A. H.; Chung, Y.-A.; and Glass, J. 2021 · 2021
Later among the works it cites.
Tera: Self-supervised learning of transformer encoder representation for speech
Liu, A. T.; Li, S.-W.; and Lee, H.-y. 2021 · 2021
Later among the works it cites.
End-to-end spoken language understanding using transformer networks and self-supervised pre-trained features
Morais, E.; Kuo, H.-K. J.; Thomas, S.; Tüske, Z.; and Kingsbury, B. 2021 · 2021
Later among the works it cites.
On the Use of External Data for Spoken Named Entity Recognition
Pasad, A.; Wu, F.; Shon, S.; Livescu, K.; and Han, K. J. 2021 · 2021
Later among the works it cites.
SLUE: New Benchmark Tasks for Spoken Language Understanding Evaluation on Natural Speech
Shon, S.; Pasad, A.; Wu, F.; Brusco, P.; Artzi, Y.; Livescu, K.; and Han, K. J. 2021 · 2021
Later among the works it cites.
Joint masked cpc and ctc training for asr
Talnikar, C.; Likhomanenko, T.; Collobert, R.; and Synnaeve, G. 2021 · 2021
Later among the works it cites.
Unispeech: Unified speech representation learning with labeled and unlabeled data
Wang, C.; Wu, Y.; Qian, Y.; Kumatani, K.; Liu, S.; Wei, F.; Zeng, M.; and Huang, X. 2021 · 2021
Later among the works it cites.
SUPERB: Speech processing Universal PERformance Benchmark
Yang, S.-w.; Chi, P.-H.; Chuang, Y.-S.; Lai, C.-I. J.; Lakhotia, K.; Lin, Y. Y.; Liu, A. T.; Shi, J.; Chang, X.; Lin, G.-T.; et al. 2021 · 2021
Later among the works it cites.