Fetching the paper…
Reading the bibliography…
Probabilistic Latent Variable Models (LVMs) provide an alternative to self-supervised learning approaches for linguistic representation learning from speech.
M. Halle and K. Stevens, “Speech recognition: A model and a program for research,”
1962
Earlier work this paper cites.
A. M. Liberman, F. S. Cooper, D. P. Shankweiler, and M. Studdert-Kennedy, “Perception of the speech code.”
1967
Earlier work this paper cites.
D. B. Paul and J. M. Baker, “The design for the wall street journal-based csr corpus,” in
1992
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in
2006
Earlier work this paper cites.
T. Hofmann, B. Schölkopf, and A. J. Smola, “Kernel methods in machine learning,”
2008
Earlier work this paper cites.
S. Goldwater, T. L. Griffiths, and M. Johnson, “A bayesian framework for word segmentation: Exploring the effects of context,”
2009
Earlier work this paper cites.
C.-y. Lee and J. Glass, “A nonparametric bayesian approach to acoustic model discovery,” in
2012
Earlier work this paper cites.
R. Ranganath, S. Gerrish, and D. M. Blei, “Black box variational inference,”
2013
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,”
2013
Earlier work this paper cites.
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
L. Ondel, L. Burget, and J. Černockỳ, “Variational inference for acoustic unit discovery,”
2016
Earlier work this paper cites.
D. Harwath, A. Torralba, and J. Glass, “Unsupervised learning of spoken language with visual context,” in
2016
Earlier work this paper cites.
M. J. Johnson, D. K. Duvenaud, A. Wiltschko, R. P. Adams, and S. R. Datta, “Composing graphical models with neural networks for structured representations and fast inference,” in
2016
Earlier work this paper cites.
B. Uria, M.-A. Côté, K. Gregor, I. Murray, and H. Larochelle, “Neural autoregressive distribution estimation,”
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
W.-N. Hsu, Y. Zhang, and J. Glass, “Unsupervised learning of disentangled and interpretable representations from sequential data,” in
2017
Cited alongside, same era.
W.-N. Hsu, Y. Zhang, and J. Glass, “Unsupervised domain adaptation for robust speech recognition via variational autoencoder-based data augmentation,” in
2019
Later among the works it cites.
A. T. Liu, S. wen Yang, P.-H. Chi, P. chun Hsu, and H. yi Lee, “Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders,” 2019
2019
Later among the works it cites.
S. Khurana, S. R. Joty, A. Ali, and J. Glass, “A factorial deep markov model for unsupervised disentangled representation learning from speech,” in
2019
Later among the works it cites.
J. Chorowski, R. J. Weiss, S. Bengio, and A. van den Oord, “Unsupervised speech representation learning using wavenet autoencoders,”
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
R. G. Krishnan, U. Shalit, and D. Sontag, “Structured inference networks for nonlinear state space models,” in
2017
Cited alongside, same era.
J. Ebbers, J. Heymann, L. Drude, T. Glarner, R. Haeb-Umbach, and B. Raj, “Hidden markov model variational autoencoder for acoustic unit discovery.” in
2017
Cited alongside, same era.
E. Dupoux, “Cognitive science in the era of artificial intelligence: A roadmap for reverse-engineering the infant language-learner,”
2018
Cited alongside, same era.
Y. Li and S. Mandt, “Disentangled sequential autoencoder,”
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Later among the works it cites.
A. Mohamed, D. Okhonko, and L. Zettlemoyer, “Transformers with convolutional context for asr,”
2019
Later among the works it cites.
L. Maaløe, M. Fraccaro, V. Liévin, and O. Winther, “Biva: A very deep hierarchy of latent variables for generative modeling,” in
2019
Later among the works it cites.
Y. Tian, D. Krishnan, and P. Isola, “Contrastive multiview coding,”
2019
Later among the works it cites.
2019
Later among the works it cites.
S. Vasquez and M. Lewis, “Melnet: A generative model for audio in the frequency domain,”
2019
Later among the works it cites.
R. Prenger, R. Valle, and B. Catanzaro, “Waveglow: A flow-based generative network for speech synthesis,” in
2019
Later among the works it cites.
D. Harwath, W.-N. Hsu, and J. Glass, “Learning hierarchical discrete linguistic units from visually-grounded speech,” in
2020
Closest in time.
W. Grathwohl, K.-C. Wang, J.-H. Jacobsen, D. Duvenaud, M. Norouzi, and K. Swersky, “Your classifier is secretly an energy based model and you should treat it like one,” in
2020
Closest in time.