Fetching the paper…
Reading the bibliography…
Speech synthesis (text to speech, TTS) and recognition (automatic speech recognition, ASR) are important speech tasks, and require a large amount of text and speech pairs for model training.
The evolutionary history of the human speech organs
Jan Wind. 1989 · 1989
Earlier work this paper cites.
Deep Voice 3: 2000-Speaker Neural Text-to-Speech. In International Conference on Learning Representations
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O. Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller. 2018 · 2000
Earlier work this paper cites.
Phonetic learning as a pathway to language: new data and native language magnet theory expanded (NLM-e)
Patricia K Kuhl, Barbara T Conboy, Sharon Coffey-Corina, Denise Padden, Maritza Rivera-Gaxiola, and Tobey Nelson. 2008 · 2008
Earlier work this paper cites.
Text-to-Speech Costs — Licensing and Pricing
Joel Harband. 2010 · 2010
Earlier work this paper cites.
Thousands of voices for HMM-based speech synthesis–Analysis and application of TTS systems built on various ASR corpora
Junichi Yamagishi, Bela Usabaev, Simon King, Oliver Watts, John Dines, Jilei Tian, Yong Guan, Rile Hu, Keiichiro Oura, Yi-Jian Wu, et al · 2010
Earlier work this paper cites.
Simons, and Charles D. Fennig (eds.).(2015). Ethnologue: Languages of the World, Dallas, Texas: SIL International
M Paul Lewis and F Gary. 2013 · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
End-to-end continuous speech recognition using attention-based recurrent nn: First results. In NIPS 2014 Workshop on Deep Learning, December 2014
Jan Chorowski, Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Effective Approaches to Attention-based Neural Machine Translation. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing . 1412–1421
Minh-Thang Luong, Hieu Pham, and Christopher D Manning. 2015 · 2015
Earlier work this paper cites.
Librispeech: an ASR corpus based on public domain audio books. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 5206–5210
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition. In Acoustics, Speech and Signal Processing (ICASSP), 2016 IEEE International Conference on . IEEE, 4960–4964
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals. 2016 · 2016
Earlier work this paper cites.
Sequence-Level Knowledge Distillation. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing . 1317–1327
Yoon Kim and Alexander M Rush. 2016 · 2016
Earlier work this paper cites.
Improving Neural Machine Translation Models with Monolingual Data. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Vol. 1. 86–96
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Comparison of grapheme-to-phoneme conversion methods on a myanmar pronunciation dictionary. In Proceedings of the 6th Workshop on South and Southeast Asian Natural Language Processing (WSSANLP2016) . 11–22
Ye Kyaw Thu, Win Pa Pa, Yoshinori Sagisaka, and Naoto Iwahashi. 2016 · 2016
Earlier work this paper cites.
Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline. In 2017 20th Conference of the Oriental Chapter of the International Coordinating Committee on Speech Databases and Speech I/O Systems and Assessment (O-COCOSDA) . IEEE, 1–5
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng. 2017 · 2017
Earlier work this paper cites.
The LJ Speech Dataset
Keith Ito. 2017 · 2017
Cited alongside, same era.
Using the Output Embedding to Improve Language Models. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers . 157–163
Ofir Press and Lior Wolf. 2017 · 2017
Cited alongside, same era.
Listening while speaking: Speech chain by deep learning. In Automatic Speech Recognition and Understanding Workshop (ASRU), 2017 IEEE . IEEE, 301–308
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura. 2017 · 2017
Cited alongside, same era.
Attention is all you need. In Advances in Neural Information Processing Systems . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Tacotron: Towards End-to-End Speech Synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al · 2017
Effectiveness of self-supervised pre-training for speech recognition
Alexei Baevski, Michael Auli, and Abdelrahman Mohamed. 2019 · 2019
Later among the works it cites.
Semi-supervised training for improving data efficiency in end-to-end speech synthesis. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6940–6944
Yu-An Chung, Yuxuan Wang, Wei-Ning Hsu, Yu Zhang, and RJ Skerry-Ryan. 2019 · 2019
Later among the works it cites.
Text-to-speech synthesis using found data for low-resource languages
Erica Lindsay Cooper. 2019 · 2019
Later among the works it cites.
Cycle-consistency training for end-to-end speech recognition. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6271–6275
Takaaki Hori, Ramon Astudillo, Tomoki Hayashi, Yu Zhang, Shinji Watanabe, and Jonathan Le Roux. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Unpaired image-to-image translation using cycle-consistent adversarial networks. In Proceedings of the IEEE international conference on computer vision . 2223–2232
Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros. 2017 · 2017
Cited alongside, same era.
Dictionary Augmented Sequence-to-Sequence Neural Network for Grapheme to Phoneme Prediction
Antoine Bruguier, Anton Bakhtin, and Dravyansh Sharma. 2018 · 2018
Cited alongside, same era.
Towards Unsupervised Automatic Speech Recognition Trained by Unaligned Speech and Text only
Yi-Chen Chen, Chia-Hao Shen, Sung-Feng Huang, and Hung-yi Lee. 2018 · 2018
Cited alongside, same era.
State-of-the-art speech recognition with sequence-to-sequence models. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 4774–4778
Chung-Cheng Chiu, Tara N Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J Weiss, Kanishka Rao, Ekaterina Gonina, et al · 2018
Cited alongside, same era.
Characteristics of Text-to-Speech and Other Corpora
Erica Cooper, Emily Li, and Julia Hirschberg. 2018 · 2018
Cited alongside, same era.
Lithuanian Speech Corpus Liepa for development of human-computer interfaces working in voice recognition and synthesis mode
Sigita Laurinčiukaitė, Laimutis Telksnys, Pijus Kasparaitis, Regina Kliukienė, and Vilma Paukštytė. 2018 · 2018
Cited alongside, same era.
Completely Unsupervised Phoneme Recognition by Adversarially Learning Mapping Relationships from Audio Embeddings
Da-Rong Liu, Kuan-Yu Chen, Hung-yi Lee, and Lin-shan Lee. 2018 · 2018
Cited alongside, same era.
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, Ming Liu, and M Zhou. 2019 · 2019
Later among the works it cites.
Towards Unsupervised Speech Recognition and Synthesis with Quantized Speech Representation Learning
Alexander H Liu, Tao Tu, Hung-yi Lee, and Lin-shan Lee. 2019 · 2019
Later among the works it cites.
SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le. 2019 · 2019
Later among the works it cites.
Speech Recognition with Augmented Synthesized Speech
Andrew Rosenberg, Yu Zhang, Bhuvana Ramabhadran, Ye Jia, Pedro Moreno, Yonghui Wu, and Zelin Wu. 2019 · 2019
Later among the works it cites.
wav2vec: Unsupervised Pre-Training for Speech Recognition
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli. 2019 · 2019
Later among the works it cites.
Token-Level Ensemble Distillation for Grapheme-to-Phoneme Conversion
Hao Sun, Xu Tan, Jun-Wei Gan, Hongzhi Liu, Sheng Zhao, Tao Qin, and Tie-Yan Liu. 2019 · 2019
Later among the works it cites.
Multilingual Neural Machine Translation with Knowledge Distillation. In International Conference on Learning Representations
Xu Tan, Yi Ren, Di He, Tao Qin, and Tie-Yan Liu. 2019 · 2019
Later among the works it cites.
Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim. 2019 · 2019
Later among the works it cites.
Unsupervised Speech Recognition via Segmental Empirical Output Distribution Matching
Chih-Kuan Yeh, Jianshu Chen, Chengzhu Yu, and Dong Yu. 2019 · 2019
Later among the works it cites.
End-to-end text-to-speech for low-resource languages by cross-lingual transfer learning
Yuan-Jui Chen, Tao Tu, Cheng-chieh Yeh, and Hung-Yi Lee. 2019a · 2079
Closest in time.