Fetching the paper…
Reading the bibliography…
Learning good representations without supervision is still an open issue in machine learning, and is particularly challenging for speech signals, which are often characterized by long sequences with a complex hierarchical structure.
S. B. Davis and P. Mermelstein, “Comparison of parametric representation for monosyllabic word recognition in continuously spoken sentences,”
1980
Earlier work this paper cites.
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, D. S. Pallett, and N. L. Dahlgren, “DARPA TIMIT Acoustic Phonetic Continuous Speech Corpus CDROM,” 1993
1993
Earlier work this paper cites.
W. S. A. Paeschke, M. Kienast, “F0-contours in emotional speech,” in
1999
Earlier work this paper cites.
V. Hozjan, Z. Kacic, A. Moreno, A. Bonafonte, and A. Nogueiras, “Interface databases: Design and collection of a multilingual emotional speech database.” in
2002
Earlier work this paper cites.
M. T. Rosenstein, Z. Marx, L. P. Kaelbling, and T. G. Dietterich, “To transfer or not to transfer,”
2005
Earlier work this paper cites.
Y. Bengio, P. Lamblin, D. Popovici, and H. Larochelle, “Greedy layer-wise training of deep networks,” in
2006
Earlier work this paper cites.
G. Hinton, S. Osindero, and Y. Teh, “A fast learning algorithm for deep belief nets,”
2006
Earlier work this paper cites.
D. Povey
2011
Earlier work this paper cites.
Y. Bengio, “Deep learning of representations for unsupervised and transfer learning,” in
2012
Earlier work this paper cites.
G. Dahl, D. Yu, L. Deng, and A. Acero, “Context-dependent pre-trained deep neural networks for large vocabulary speech recognition,”
2012
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” in
2014
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in
2014
Earlier work this paper cites.
J. Chung, Ç. Gülçehre, K. Cho, and Y. Bengio, “Empirical evaluation of gated recurrent neural networks on sequence modeling,” in
2014
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,”
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,” in
2015
Cited alongside, same era.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in
2015
Cited alongside, same era.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in
2015
Cited alongside, same era.
M. Ravanelli, L. Cristoforetti, R. Gretter, M. Pellin, A. Sosi, and M. Omologo, “The DIRHA-ENGLISH corpus and related tasks for distant-speech recognition in domestic environments,” in
2015
Cited alongside, same era.
I. Misra, C. L. Zitnick, and M. Hebert, “Shuffle and learn: Unsupervised learning using temporal order verification,” in
2016
Cited alongside, same era.
A. Jansen, M. Plakal, R. Pandya, D. P. W. Ellis, S. Hershey, J. Liu, R. C. Moore, and R. A. Saurous, “Unsupervised learning of semantic audio representations,” in
2018
Later among the works it cites.
A. van den Oord, Y. Li, and O. Vinyals, “Representation learning with contrastive predictive coding,”
2018
Later among the works it cites.
M. Ravanelli and Y. Bengio, “Learning speaker representations with mutual information,”
2018
Later among the works it cites.
D. Hjelm, A. Fedorov, S. Lavoie-Marchildon, K. Grewal, A. Trischler, and Y. Bengio, “Learning deep representations by mutual information estimation and maximization,”
2018
Later among the works it cites.
M. Ravanelli and Y. Bengio, “Speaker recognition from raw waveform with SincNet,”
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in
2016
Cited alongside, same era.
C. Veaux, J. Yamagishi, and K. MacDonald, “CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit,” 2016
2016
Cited alongside, same era.
S. Pascual, A. Bonafonte, and J. Serrà, “Segan: Speech enhancement generative adversarial network,” in
2017
Cited alongside, same era.
M. Neumann and N. T. Vu, “Attentive convolutional neural network based speech emotion recognition: A study on the impact of input features, signal length, and acted speech,” in
2017
Cited alongside, same era.
A. van den Oord, O. Vinyals, and K. Kavukcuoglu, “Neural discrete representation learning,” in
2017
Cited alongside, same era.
S. Chang, Y. Zhang, W. Han, M. Yu, X. Guo, W. Tan, X. Cui, M. J. Witbrock, M. Hasegawa-Johnson, and T. S. Huang, “Dilated recurrent neural networks,” in
2017
Cited alongside, same era.
J. Serrà, S. Pascual, and A. Karatzoglou, “Towards a universal neural network encoder for time series,” in
2018
Cited alongside, same era.
M. Ravanelli and Y.Bengio, “Interpretable convolutional filters with SincNet,”
2018
Later among the works it cites.
M. Ravanelli and M. Omologo, “Automatic context window composition for distant speech recognition,”
2018
Later among the works it cites.
M. Ravanelli, P. Brakel, M. Omologo, and Y. Bengio, “Light gated recurrent units for speech recognition,”
2018
Later among the works it cites.
J. Michálek and J. Vanek, “A survey of recent DNN architectures on the TIMIT phone recognition task,” in
2018
Later among the works it cites.
2019
Closest in time.
B. McFee
2019
Closest in time.
R. Yamamoto, J. Felipe, and M. Blaauw, “r9y9/pysptk: 0.1.14,” Jan. 2019. [Online]. Available:
2019
Closest in time.
M. Ravanelli, T. Parcollet, and Y. Bengio, “The PyTorch-Kaldi Speech Recognition Toolkit,” in
2019
Closest in time.
J. Wang, K. Wang, M. Law, F. Rudzicz, and M. Brudno, “Centroid-based deep metric learning for speaker recognition,” in
2019
Closest in time.
C. Doersch and A. Zisserman, “Multi-task self-supervised visual learning,” in
2079
Closest in time.