Fetching the paper…
Reading the bibliography…
We consider the task of unsupervised extraction of meaningful latent representations of speech by applying autoencoding neural networks to speech waveforms.
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning representations by back-propagating errors,” Nature , vol. 323, no. 6088, 1986
1986
Earlier work this paper cites.
B. T. Polyak and A. B. Juditsky, “Acceleration of stochastic approximation by averaging,” SIAM Journal on Control and Optimization , vol. 30, no. 4, pp. 838–855, 1992
1992
Earlier work this paper cites.
B. A. Olshausen and D. J. Field, “Emergence of simple-cell receptive field properties by learning a sparse code for natural images,” Nature , vol. 381, no. 6583, p. 607, 1996
1996
Earlier work this paper cites.
X. Wang and C.-C. J. Kuo, “An 800 bps VQ-based LPC voice coder,” Journal of the Acoustical Society of America , vol. 103, no. 5, 1998
1998
Earlier work this paper cites.
D. D. Lee and H. S. Seung, “Learning the parts of objects by non-negative matrix factorization,” Nature , vol. 401, no. 6755, p. 788, 1999
1999
Earlier work this paper cites.
L. Wiskott and T. J. Sejnowski, “Slow feature analysis: Unsupervised learning of invariances,” Neural Computation , vol. 14, no. 4, 2002
2002
Earlier work this paper cites.
G. E. Hinton and R. R. Salakhutdinov, “Reducing the dimensionality of data with neural networks,” Science , vol. 313, no. 5786, 2006
2006
Earlier work this paper cites.
C. M. Bishop, “Continuous latent variables,” in Pattern Recognition and Machine Learning . Springer, 2006, ch. 12
2006
Earlier work this paper cites.
P. Vincent, H. Larochelle, Y. Bengio, and P.-A. Manzagol, “Extracting and composing robust features with denoising autoencoders,” in Proc. International Conference on Machine Learning , 2008
2008
Earlier work this paper cites.
H. Lee, C. Ekanadham, and A. Ng, “Sparse deep belief net model for visual area V2,” in Advances in Neural Information Processing Systems , 2008
2008
Earlier work this paper cites.
A. S. Park and J. R. Glass, “Unsupervised Pattern Discovery in Speech,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 16, no. 1, pp. 186–197, Jan. 2008
2008
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A Large-Scale Hierarchical Image Database,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition , 2009
2009
Earlier work this paper cites.
D. Jurafsky and J. H. Martin, Speech and Language Processing (2nd Edition) . Upper Saddle River, NJ, USA: Prentice-Hall, Inc., 2009
2009
Earlier work this paper cites.
N. Jaitly and G. Hinton, “Learning a better representation of speech soundwaves using restricted Boltzmann machines,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2011
2011
Earlier work this paper cites.
D. Yu and M. L. Seltzer, “Improved bottleneck features using pretrained deep neural networks,” in Proc. Interspeech , 2011
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely, “The Kaldi speech recognition toolkit,” in Proc. Automatic Speech Recognition and Understanding Workshop (ASRU) , 2011
2011
Earlier work this paper cites.
A. Jansen and B. Van Durme, “Efficient spoken term discovery using randomized algorithms,” in Proc. Automatic Speech Recognition and Understanding Workshop (ASRU) , 2011, pp. 401–406
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
P. Swietojanski, A. Ghoshal, and S. Renals, “Unsupervised cross-lingual knowledge transfer in DNN-based LVCSR,” in Proc. Spoken Language Technology Workshop (SLT) , 2012, pp. 246–251
2012
Earlier work this paper cites.
K. Veselỳ, M. Karafiát, F. Grézl, M. Janda, and E. Egorova, “The language-independent bottleneck features,” in Proc. Spoken Language Technology Workshop (SLT) , 2012, pp. 336–341
2012
Earlier work this paper cites.
C.-y. Lee and J. Glass, “A Nonparametric Bayesian Approach to Acoustic Model Discovery,” in Proc. 50th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Jul. 2012, pp. 40–49
2012
Earlier work this paper cites.
A. Graves, A.-r. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2013, pp. 6645–6649
2013
Earlier work this paper cites.
S. Thomas, M. L. Seltzer, K. Church, and H. Hermansky, “Deep neural network features and semi-supervised training for low resource speech recognition,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2013, pp. 6704–6708
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
T. Schatz, V. Peddinti, F. Bach, A. Jansen, H. Hermansky, and E. Dupoux, “Evaluating speech features with the minimal-pair ABX task: Analysis of the classical MFC/PLP pipeline,” in Proc. Interspeech , 2013, pp. 1–5
2013
Earlier work this paper cites.
M. D. Zeiler and R. Fergus, “Visualizing and understanding convolutional networks,” in European Conference on Computer Vision , 2014
2014
Earlier work this paper cites.
S. Dieleman and B. Schrauwen, “End-to-end learning for music audio,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2014, pp. 6964–6968
2014
Earlier work this paper cites.
Z. Tüske, P. Golik, R. Schlüter, and H. Ney, “Acoustic modeling with deep neural networks using raw time signal for LVCSR,” in Proc. Interspeech , 2014
2014
Cited alongside, same era.
J. Yosinski, J. Clune, Y. Bengio, and H. Lipson, “How transferable are features in deep neural networks?” in Advances in Neural Information Processing Systems , 2014, pp. 3320–3328
2014
Cited alongside, same era.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” in Proc. International Conference on Learning Representations , 2014
2014
Cited alongside, same era.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A simple way to prevent neural networks from overfitting,” Journal of Machine Learning Research , vol. 15, no. 1, 2014
2014
Cited alongside, same era.
A. Conneau, D. Kiela, H. Schwenk, L. Barrault, and A. Bordes, “Supervised learning of universal sentence representations from natural language inference data,” in Proc. Conference on Empirical Methods in Natural Language Processing (EMNLP) , September 2017, pp. 670–680
2017
Later among the works it cites.
I. Higgins, L. Matthey, A. Pal, C. Burgess, X. Glorot, M. Botvinick, S. Mohamed, and A. Lerchner, “Beta-VAE: Learning basic visual concepts with a constrained variational framework,” in Proc. International Conference on Learning Representations , 2017
2017
Later among the works it cites.
A. A. Alemi, I. Fischer, J. V. Dillon, and K. Murphy, “Deep variational information bottleneck,” in Proc. International Conference on Learning Representations , 2017
2017
Later among the works it cites.
Y. N. Dauphin, A. Fan, M. Auli, and D. Grangier, “Language modeling with gated convolutional networks,” in Proc. International Conference on Machine Learning , 2017
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition , 2015
2015
Cited alongside, same era.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in Proc. International Conference on Learning Representations , 2015
2015
Cited alongside, same era.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “LibriSpeech: an ASR corpus based on public domain audio books,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015
2015
Cited alongside, same era.
D. Palaz, M. Magima Doss, and R. Collobert, “Analysis of CNN-based speech recognition system using raw speech as input,” in Proc. Interspeech , 2015
2015
Cited alongside, same era.
T. N. Sainath, R. J. Weiss, A. Senior, K. W. Wilson, and O. Vinyals, “Learning the speech front-end with raw waveform CLDNNs,” in Proc. Interspeech , 2015
2015
Cited alongside, same era.
S. R. Bowman, G. Angeli, C. Potts, and C. Manning, “A large annotated corpus for learning natural language inference,” in Proc. Conference on Empirical Methods in Natural Language Processing , 2015
2015
Cited alongside, same era.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proc. International Conference on Learning Representations , 2015
2015
Cited alongside, same era.
I. Gulrajani, K. Kumar, F. Ahmed, A. A. Taïga, F. Visin, D. Vázquez, and A. Courville, “PixelVAE: A latent variable model for natural images,” in Proc. International Conference on Learning Representations , 2017
2017
Later among the works it cites.
D. Krueger, T. Maharaj, J. Kramár, M. Pezeshki, N. Ballas, N. R. Ke, A. Goyal, Y. Bengio, A. Courville, and C. Pal, “Zoneout: Regularizing RNNs by randomly preserving hidden activations,” in Proc. International Conference on Learning Representations , 2017
2017
Later among the works it cites.
H. Chen, C.-C. Leung, L. Xie, B. Ma, and H. Li, “Multilingual bottle-neck feature learning from untranscribed speech,” in Proc. Automatic Speech Recognition and Understanding Workshop (ASRU) , 2017
2017
Later among the works it cites.
T. Ansari, R. Kumar, S. Singh, and S. Ganapathy, “Deep learning methods for unsupervised acoustic modeling—leap submission to zerospeech challenge 2017,” in Proc. Automatic Speech Recognition and Understanding Workshop (ASRU) , 2017, pp. 754–761
2017
Later among the works it cites.
Y. Yuan, C. C. Leung, L. Xie, H. Chen, B. Ma, and H. Li, “Extracting bottleneck features and word-like pairs from untranscribed speech for feature representation,” in Proc. Automatic Speech Recognition and Understanding Workshop (ASRU) , Dec 2017, pp. 734–739
2017
Later among the works it cites.
J. Engel, C. Resnick, A. Roberts, S. Dieleman, M. Norouzi, D. Eck, and K. Simonyan, “Neural audio synthesis of musical notes with wavenet autoencoders,” in Proc. International Conference on Machine Learning , 2017, pp. 1068–1077
2017
Later among the works it cites.
W.-N. Hsu, Y. Zhang, and J. Glass, “Unsupervised learning of disentangled and interpretable representations from sequential data,” in Advances in Neural Information Processing Systems , 2017, pp. 1876–1887
2017
Later among the works it cites.
J. Ebbers, J. Heymann, L. Drude, T. Glarner, R. Haeb-Umbach, and B. Raj, “Hidden Markov Model Variational Autoencoder for Acoustic Unit Discovery,” in Proc. Interspeech , Aug. 2017, pp. 488–492
2017
Later among the works it cites.
H. Kamper, A. Jansen, and S. Goldwater, “A segmental framework for fully-unsupervised large-vocabulary speech recognition,” Computer Speech & Language , vol. 46, pp. 154–174, 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
C.-C.Chiu, T. N. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R. J. Weiss, K. Rao, K. Gonina, N. Jaitly, B. Li, J. Chorowski, and M. Bacchiani, “State-of-the-art speech recognition with sequence-to-sequence models,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018
2018
Later among the works it cites.
A. W. Yu, D. Dohan, M.-T. Luong, R. Zhao, K. Chen, M. Norouzi, and Q. V. Le, “QANet: Combining local convolution with global self-attention for reading comprehension,” in Proc. International Conference on Learning Representations , 2018
2018
Later among the works it cites.
J. Chorowski, R. J. Weiss, R. A. Saurous, and S. Bengio, “On using backpropagation for speech texture generation and voice conversion,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Apr. 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
D. Moyer, S. Gao, R. Brekelmans, A. Galstyan, and G. Ver Steeg, “Invariant Representations without Adversarial Training,” in Advances in Neural Information Processing Systems 31 , 2018, pp. 9084–9093
2018
Later among the works it cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. J. Skerry-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu, “Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018
2018
Later among the works it cites.
Y. Li and S. Mandt, “Disentangled sequential autoencoder,” in Proc. International Conference on Machine Learning , 2018
2018
Later among the works it cites.
T. Glarner, P. Hanebrink, J. Ebbers, and R. Haeb-Umbach, “Full Bayesian Hidden Markov Model Variational Autoencoder for Acoustic Unit Discovery,” in Proc. Interspeech , Sep. 2018, pp. 2688–2692
2018
Later among the works it cites.
G. Lample, A. Conneau, L. Denoyer, and M. Ranzato, “Unsupervised Machine Translation Using Monolingual Corpora Only,” in Proc. International Conference on Learning Representations , 2018
2018
Later among the works it cites.
Y.-A. Chung, W.-H. Weng, S. Tong, and J. Glass, “Unsupervised cross-modal alignment of speech and text embedding spaces,” Advances in Neural Information Processing Systems , 2018
2018
Later among the works it cites.
W.-N. Hsu, Y. Zhang, R. J. Weiss, H. Zen, Y. Wu, Y. Wang, Y. Cao, Y. Jia, Z. Chen, J. Shen, P. Nguyen, and R. Pang, “Hierarchical generative modeling for controllable speech synthesis,” in Proc. International Conference on Learning Representations , 2019
2019
Closest in time.
2019
Closest in time.