Fetching the paper…
Reading the bibliography…
Deep generative models have achieved great success in unsupervised learning with the ability to capture complex nonlinear relationships between latent generating factors and observations.
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, and D. S. Pallett, “DARPA TIMIT acoustic-phonetic continous speech corpus CD-ROM. NIST speech disc 1-1.1,”
1993
Earlier work this paper cites.
A. K. Halberstadt, “Heterogeneous acoustic measurements and multiple classifiers for speech recognition,” Ph.D. dissertation, Massachusetts Institute of Technology, 1999
1999
Earlier work this paper cites.
D. Pearce, “Aurora working group: Dsr front end lvcsr evaluation au/384/02,” Ph.D. dissertation, Mississippi State University, 2002
2002
Earlier work this paper cites.
J. Carletta, “Unleashing the killer corpus: experiences in creating the multi-everything ami meeting corpus,”
2007
Earlier work this paper cites.
J. Garofalo, D. Graff, D. Paul, and D. Pallett, “Csr-i (wsj0) complete,”
2007
Earlier work this paper cites.
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,”
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz
2011
Cited alongside, same era.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,”
2013
Cited alongside, same era.
L. Van Der Maaten, “Accelerating t-sne using tree-based algorithms.”
2014
Cited alongside, same era.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,”
2014
Cited alongside, same era.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in
2015
Cited alongside, same era.
J. F. Drexler, “Deep unsupervised learning from speech,” 2016
2016
Later among the works it cites.
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard
2016
Later among the works it cites.
W.-N. Hsu, Y. Zhang, and J. Glass, “Learning latent representations for speech generation and transformation,” in
2017
Later among the works it cites.
——, “Unsupervised learning of disentangled and interpretable representations from sequential data,” in
2017
Later among the works it cites.
W.-N. Hsu, Y. Zhang, and J. Glass, “Unsupervised domain adaptation for robust speech recognition via variational autoencoder-based data augmentation,” in
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Tan and K. C. Sim, “Learning utterance-level normalisation using variational autoencoders for robust automatic speech recognition,” in
2016
Cited alongside, same era.
2018
Closest in time.