Fetching the paper…
Reading the bibliography…
Text to speech (TTS) and automatic speech recognition (ASR) are two dual tasks in speech processing and both achieve impressive performance thanks to the recent advance in deep learning and large amount of aligned speech and text data.
Factors governing the intelligibility of speech sounds
French, N. R. and Steinberg, J. C · 1947
Earlier work this paper cites.
Signal estimation from modified short-time fourier transform
Griffin, D. and Lim, J · 1984
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Vincent, P., Larochelle, H., Bengio, Y., and Manzagol, P.-A · 2008
Earlier work this paper cites.
End-to-end continuous speech recognition using attention-based recurrent nn: First results
Chorowski, J., Bahdanau, D., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Bengio, S., Vinyals, O., Jaitly, N., and Shazeer, N · 2015
Earlier work this paper cites.
Attention-based models for speech recognition
Chorowski, J. K., Bahdanau, D., Serdyuk, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
Chan, W., Jaitly, N., Le, Q., and Vinyals, O · 2016
Earlier work this paper cites.
Multilingual techniques for low resource automatic speech recognition
Chuangsuwanich, E · 2016
Earlier work this paper cites.
Dual learning for machine translation
He, D., Xia, Y., Qin, T., Wang, L., Yu, N., Liu, T.-Y., and Ma, W.-Y · 2016
Earlier work this paper cites.
Improving neural machine translation models with monolingual data
Sennrich, R., Haddow, B., and Birch, A · 2016
Earlier work this paper cites.
Minimum risk training for neural machine translation
Shen, S., Cheng, Y., He, Z., He, W., Wu, H., Sun, M., and Liu, Y · 2016
Earlier work this paper cites.
Deep voice: Real-time neural text-to-speech
Arik, S. O., Chrzanowski, M., Coates, A., Diamos, G., Gibiansky, A., Kang, Y., Li, X., Miller, J., Ng, A., Raiman, J., et al · 2017
Earlier work this paper cites.
Unsupervised neural machine translation
Artetxe, M., Labaka, G., Agirre, E., and Cho, K · 2017
Earlier work this paper cites.
The lj speech dataset
Ito, K · 2017
Cited alongside, same era.
Unsupervised machine translation using monolingual corpora only
Lample, G., Conneau, A., Denoyer, L., and Ranzato, M · 2017
Cited alongside, same era.
Listening while speaking: Speech chain by deep learning
Tjandra, A., Sakti, S., and Nakamura, S · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Tacotron: Towards end-to-end speech synthesis
Wang, Y., Skerry-Ryan, R., Stanton, D., Wu, Y., Weiss, R. J., Jaitly, N., Yang, Z., Xiao, Y., Chen, Z., Bengio, S., et al · 2017
Cited alongside, same era.
Neural voice cloning with a few samples
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions
Shen, J., Pang, R., Weiss, R. J., Schuster, M., Jaitly, N., Yang, Z., Chen, Z., Zhang, Y., Wang, Y., Skerrv-Ryan, R., et al · 2018
Later among the works it cites.
Machine speech chain with one-shot speaker adaptation
Tjandra, A., Sakti, S., and Nakamura, S · 2018
Later among the works it cites.
Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis
Wang, Y., Stanton, D., Zhang, Y., Skerry-Ryan, R. J., Battenberg, E., Shor, J., Xiao, Y., Jia, Y., Ren, F., and Saurous, R. A · 2018
Later among the works it cites.
Beyond error propagation in neural machine translation: Characteristics of language also matter
Wu, L., Tan, X., He, D., Tian, F., Qin, T., Lai, J., and Liu, T.-Y · 2018
Later among the works it cites.
The microsoft 2017 conversational speech recognition system
Xiong, W., Wu, L., Alleva, F., Droppo, J., Huang, X., and Stolcke, A · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Arik, S. Ö., Chen, J., Peng, K., Ping, W., and Zhou, Y · 2018
Cited alongside, same era.
State-of-the-art speech recognition with sequence-to-sequence models
Chiu, C.-C., Sainath, T. N., Wu, Y., Prabhavalkar, R., Nguyen, P., Chen, Z., Kannan, A., Weiss, R. J., Rao, K., Gonina, E., et al · 2018
Cited alongside, same era.
Sequence-based multi-lingual low resource speech recognition
Dalmia, S., Sanabria, R., Metze, F., and Black, A. W · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Transfer learning from speaker verification to multispeaker text-to-speech synthesis
Jia, Y., Zhang, Y., Weiss, R. J., Wang, Q., Shen, J., Ren, F., Chen, Z., Nguyen, P., Pang, R., Lopez-Moreno, I., and Wu, Y · 2018
Cited alongside, same era.
Phrase-based & neural unsupervised machine translation
Lample, G., Ott, M., Conneau, A., Denoyer, L., and Ranzato, M · 2018
Cited alongside, same era.
Close to human quality tts with transformer
Li, N., Liu, S., Liu, Y., Zhao, S., Liu, M., and Zhou, M · 2018
Cited alongside, same era.
Yang, X., Li, J., and Zhou, X · 2018
Later among the works it cites.
Deep-fsmn for large vocabulary continuous speech recognition
Zhang, S., Lei, M., Yan, Z., and Dai, L · 2018
Later among the works it cites.
Multilingual end-to-end speech recognition with a single transformer on low-resource languages
Zhou, S., Xu, S., and Xu, B · 2018
Later among the works it cites.
Sample efficient adaptive text-to-speech
Chen, Y., Assael, Y., Shillingford, B., Budden, D., Reed, S., Zen, H., Wang, Q., Cobo, L. C., Trask, A., Laurie, B., Gulcehre, C., van den Oord, A., Vinyals, O., and de Freitas, N · 2019
Closest in time.
Clarinet: Parallel wave generation in end-to-end text-to-speech
Ping, W., Peng, K., and Chen, J · 2019
Closest in time.
Fastspeech: Fast, robust and controllable text to speech
Ren, Y., Ruan, Y., Tan, X., Qin, T., Zhao, S., Zhao, Z., and Liu, T.-Y · 2019
Closest in time.
Mass: Masked sequence to sequence pre-training for language generation
Song, K., Tan, X., Qin, T., Lu, J., and Liu, T.-Y · 2019
Closest in time.
Token-level ensemble distillation for grapheme-to-phoneme conversion
Sun, H., Tan, X., Gan, J.-W., Liu, H., Zhao, S., Qin, T., and Liu, T.-Y · 2019
Closest in time.
Unsupervised speech recognition via segmental empirical output distribution matching
Yeh, C.-K., Chen, J., Yu, C., and Yu, D · 2019
Closest in time.