“Efficient estimation of word representations in vector space,”
Original
T. Mikolov, K. Chen, G. Corrado, and J. Dean, · 2013
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Original
D. P Kingma and J. Ba, · 2014
Earlier work this paper cites.
“Unsupervised learning of visual representations using videos,”
X. Wang and A. Gupta, · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“Musan: A music, speech, and noise corpus,”
Original
D. Snyder, G. Chen, and D. Povey, · 2015
Earlier work this paper cites.
“Audio word2vec: Unsupervised learning of audio segment representations using sequence-to-sequence autoencoder,”
Original
Y.A. Chung, C.C. Wu, C.H. Shen, H.Y Lee, and L.S. Lee, · 2016
Earlier work this paper cites.
“Audio enhancing with dnn autoencoder for speaker recognition,”
O. Plchot, L. Burget, H. Aronowitz, and P. Matejka, · 2016
Earlier work this paper cites.
“Layer normalization,”
Original
J.L Ba, J.R. Kiros, and G.E. Hinton, · 2016
Earlier work this paper cites.
“Unsupervised feature learning for audio analysis,”
Original
M. Meyer, J. Beutel, and L. Thiele, · 2017
Earlier work this paper cites.
“Google’s next-generation real-time unit-selection synthesizer using sequence-to-sequence lstm-based autoencoders.,”
V. Wan, Y. Agiomyrgiannakis, H. Silen, and J. Vit, · 2017
Earlier work this paper cites.
“Audio set: An ontology and human-labeled dataset for audio events,”
J. F. Gemmeke, D.P.W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R.C. Moore, M. Plakal, and M. Ritter, · 2017
Earlier work this paper cites.
“Voxceleb: a large-scale speaker identification dataset,”
Original
A. Nagrani, J.S. Chung, and A. Zisserman, · 2017
Earlier work this paper cites.