Fetching the paper…
Reading the bibliography…
In this paper we propose a Sequential Representation Quantization AutoEncoder (SeqRQ-AE) to learn from primarily unpaired audio data and produce sequences of representations very close to phoneme sequences of speech utterances.
“Continuous speech recognition by statistical methods,”
Frederick Jelinek, · 1976
Earlier work this paper cites.
“Signal estimation from modified short-time fourier transform,”
Daniel Griffin and Jae Lim, · 1984
Earlier work this paper cites.
Handbook of the International Phonetic Association: A Guide to the Use of the International Phonetic Alphabet
International Phonetic Association, C.U. Press, and International Phonetic Association Staff, · 1999
Earlier work this paper cites.
“Identifying speakers in children’s stories for speech synthesis,”
Jason Y Zhang, Alan W Black, and Richard Sproat, · 2003
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Visualizing data using t-sne,”
Laurens van der Maaten and Geoffrey Hinton, · 2008
Earlier work this paper cites.
“Statistical parametric speech synthesis,”
Heiga Zen, Keiichi Tokuda, and Alan W Black, · 2009
Earlier work this paper cites.
“Estimating or propagating gradients through stochastic neurons for conditional computation,”
Yoshua Bengio, Nicholas Léonard, and Aaron Courville, · 2013
Earlier work this paper cites.
“Audio word2vec: Unsupervised learning of audio segment representations using sequence-to-sequence autoencoder,”
Yu-An Chung, Chao-Chung Wu, Chia-Hao Shen, Hung-Yi Lee, and Lin-Shan Lee, · 2016
Earlier work this paper cites.
“Neural discrete representation learning,”
Aaron van den Oord, Oriol Vinyals, et al., · 2017
Cited alongside, same era.
“Listening while speaking: Speech chain by deep learning,”
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura, · 2017
Cited alongside, same era.
“The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
Keith Ito, · 2017
Cited alongside, same era.
“Tacotron: Towards end-to-end speech synthesis,”
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al., · 2017
Cited alongside, same era.
“Representation learning with contrastive predictive coding,”
Aaron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Cited alongside, same era.
“wav2vec: Unsupervised Pre-Training for Speech Recognition,”
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli, · 2019
Closest in time.
“An unsupervised autoregressive model for speech representation learning,”
Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James Glass, · 2019
Closest in time.
“Unsupervised speech representation learning using wavenet autoencoders,”
Jan Chorowski, Ron J Weiss, Samy Bengio, and Aäron van den Oord, · 2019
Closest in time.
“vq-wav2vec: Self-supervised learning of discrete speech representations,”
Alexei Baevski, Steffen Schneider, and Michael Auli, · 2019
Closest in time.
“End-to-end feedback loss in speech chain framework via straight-through estimator,”
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura, · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Segmental audio word2vec: Representing utterances as sequences of vectors with applications in spoken term detection,”
Yu-Hsuan Wang, Hung-yi Lee, and Lin-shan Lee, · 2018
Cited alongside, same era.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al., · 2018
Cited alongside, same era.
“Phonological mappings for english, french, german and portuguese,”
Sibo Tong and Philip N. Garner, · 2018
Cited alongside, same era.
Closest in time.
“Almost unsupervised text to speech and automatic speech recognition,”
Yi Ren, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2019
Closest in time.
“g2pe,” https://github.com/Kyubyong/g2p , 2019
Kyubyong Park and Jongseok Kim, · 2019
Closest in time.
“Sequence to sequence neural speech synthesis with prosody modification capabilities,”
Slava Shechtman and Alex Sorin, · 2019
Closest in time.