Fetching the paper…
Reading the bibliography…
This paper proposes a novel unsupervised autoregressive neural model for learning generic speech representations.
M. Schroeder and B. Atal, “Code-excided linear prediction (CELP): high-quality speech at very low bit rates,” in
1985
Earlier work this paper cites.
D. Paul and J. Baker, “The design for the wall street journal-based csr corpus,” in
1992
Earlier work this paper cites.
R. Caruana, “Multitask learning,”
1997
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,”
1997
Earlier work this paper cites.
N. Tishby, F. Pereira, and W. Bialek, “The information bottleneck method,”
1999
Earlier work this paper cites.
P. Vincent, H. Larochelle, I. Lajoie, Y. Bengio, and P.-A. Manzagol, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion,”
2010
Earlier work this paper cites.
T. Mikolov, M. Karafiát, L. Burget, J. Černockỳ, and S. Khudanpur, “Recurrent neural network based language model,” in
2010
Earlier work this paper cites.
N. Dehak, P. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,”
2011
Earlier work this paper cites.
T. Schatz, V. Peddinti, F. Bach, A. Jansen, H. Hermansky, and E. Dupoux, “Evaluating speech features with the minimal-pair ABX task: Analysis of the classical MFC/PLP pipeline,” in
2013
Earlier work this paper cites.
X. Wang and A. Gupta, “Unsupervised learning of visual representations using videos,” in
2015
Earlier work this paper cites.
C. Doersch, A. Gupta, and A. Efros, “Unsupervised visual representation learning by context prediction,” in
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in
2015
Cited alongside, same era.
S. Settle and K. Livescu, “Discriminative acoustic word embeddings: recurrent neural network-based approaches,” in
2016
Cited alongside, same era.
H. Kamper, W. Wang, and K. Livescu, “Deep convolutional acoustic word embeddings using word-pair side information,” in
2016
Cited alongside, same era.
Y.-A. Chung, C.-C. Wu, C.-H. Shen, H.-Y. Lee, and L.-S. Lee, “Audio word2vec: Unsupervised learning of audio segment representations using sequence-to-sequence autoencoder,” in
2016
Cited alongside, same era.
2016
Cited alongside, same era.
B. Milde and C. Biemann, “Unspeech: Unsupervised speech context embeddings,” in
2018
Later among the works it cites.
M. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer, “Deep contextualized word representations,” in
2018
Later among the works it cites.
J. Howard and S. Ruder, “Universal language model fine-tuning for text classification,” in
2018
Later among the works it cites.
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving language understanding by generative pre-training,” OpenAI, Tech. Rep., 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Cited alongside, same era.
2016
Cited alongside, same era.
G. Larsson, M. Maire, and G. Shakhnarovich, “Colorization as a proxy task for visual understanding,” in
2017
Cited alongside, same era.
W.-N. Hsu, Y. Zhang, and J. Glass, “Unsupervised learning of disentangled and interpretable representations from sequential data,” in
2017
Cited alongside, same era.
——, “Learning latent representations for speech generation and transformation,” in
2017
Cited alongside, same era.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. Saurous, “Tacotron: Towards end-to-end speech synthesis,” in
2017
Cited alongside, same era.
Y.-A. Chung and J. Glass, “Speech2vec: A sequence-to-sequence framework for learning word embeddings from speech,” in
2018
Cited alongside, same era.
2018
Later among the works it cites.
2018
Later among the works it cites.
2019
Closest in time.
2019
Closest in time.
Y.-A. Chung, Y. Wang, W.-N. Hsu, Y. Zhang, and R. Skerry-Ryan, “Semi-supervised training for improving data efficiency in end-to-end speech synthesis,” in
2019
Closest in time.
H. Tang and J. Glass, “On training recurrent networks with truncated backpropagation through time in speech recognition,” in
2019
Closest in time.