Fetching the paper…
Reading the bibliography…
Machine Speech Chain, which integrates both end-to-end (E2E) automatic speech recognition (ASR) and text-to-speech (TTS) into one circle for joint training, has been proven to be effective in data augmentation by leveraging large amounts of unpaired data.
J. Neto, L. Almeida, M. Hochberg, C. Martins, L. Nunes, S. Renals, and T. Robinson, “Speaker-adaptation for hybrid hmm-ann continuous speech recognition system,” 1995
1995
Earlier work this paper cites.
J. Cohen, T. Kamm, and A. G. Andreou, “Vocal tract normalization in speech recognition: Compensating for systematic speaker variability,” The Journal of the Acoustical Society of America , vol. 97, no. 5, pp. 3246–3247, 1995
1995
Earlier work this paper cites.
M. J. Gales, “Maximum likelihood linear transformations for hmm-based speech recognition,” Computer speech & language , vol. 12, no. 2, pp. 75–98, 1998
1998
Earlier work this paper cites.
M. Chu, C. Li, H. Peng, and E. Chang, “Domain adaptation for tts systems,” in 2002 IEEE International Conference on Acoustics, Speech, and Signal Processing , vol. 1. IEEE, 2002, pp. I–453
2002
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proceedings of the 23rd international conference on Machine learning , 2006, pp. 369–376
2006
Earlier work this paper cites.
——, “Enhancing Automatic Speech Recognition for Romanian by Using Machine Translated and Web-based Text Corpora,” in SPECOM’2011 , Kazan, Russia, 2011, pp. x–x. [Online]. Available: https://hal.archives-ouvertes.fr/hal-00959159
2011
Earlier work this paper cites.
F. Seide, G. Li, X. Chen, and D. Yu, “Feature engineering in context-dependent deep neural networks for conversational speech transcription,” in 2011 IEEE Workshop on Automatic Speech Recognition & Understanding . IEEE, 2011, pp. 24–29
2011
Earlier work this paper cites.
2012
Earlier work this paper cites.
H. Cucu, L. Besacier, C. Burileanu, and A. Buzo, “Asr domain adaptation methods for low-resourced languages: Application to romanian language,” in 2012 Proceedings of the 20th European Signal Processing Conference (EUSIPCO) . IEEE, 2012, pp. 1648–1652
2012
Earlier work this paper cites.
A. Rousseau, P. Deléglise, and Y. Esteve, “Ted-lium: an automatic speech recognition dedicated corpus.” in LREC , 2012, pp. 125–129
2012
Earlier work this paper cites.
M. D. Zeiler, “Adadelta: an adaptive learning rate method,” arXiv preprint arXiv:1212.5701 , 2012
2012
Earlier work this paper cites.
G. Saon, H. Soltau, D. Nahamoo, and M. Picheny, “Speaker adaptation of neural network acoustic models using i-vectors,” in 2013 IEEE Workshop on Automatic Speech Recognition and Understanding . IEEE, 2013, pp. 55–59
2013
Earlier work this paper cites.
P. Swietojanski and S. Renals, “Learning hidden unit contributions for unsupervised speaker adaptation of neural network acoustic models,” in 2014 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2014, pp. 171–176
2014
Cited alongside, same era.
2014
Cited alongside, same era.
2015
Cited alongside, same era.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2015, pp. 5206–5210
2015
Cited alongside, same era.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust dnn embeddings for speaker recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5329–5333
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in Advances in neural information processing systems , 2015, pp. 577–585
2015
Cited alongside, same era.
A. Tjandra, S. Sakti, and S. Nakamura, “Listening while speaking: Speech chain by deep learning,” in 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2017, pp. 301–308
2017
Cited alongside, same era.
2017
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan et al. , “Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 4779–4783
2018
Cited alongside, same era.
M. Delcroix, S. Watanabe, A. Ogawa, S. Karita, and T. Nakatani, “Auxiliary feature based adaptation of end-to-end asr systems.” in Interspeech , 2018, pp. 2444–2448
2018
Cited alongside, same era.
2018
Cited alongside, same era.
D.-R. Liu, C.-Y. Yang, S.-L. Wu, and H.-Y. Lee, “Improving unsupervised style transfer in end-to-end speech synthesis with end-to-end speech recognition,” in 2018 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2018, pp. 640–647
2018
Cited alongside, same era.
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. Enrique Yalta Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, and T. Ochiai, “Espnet: End-to-end speech processing toolkit,” in Interspeech , 2018, pp. 2207–2211. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2018-1456
2018
Cited alongside, same era.
Later among the works it cites.
2019
Later among the works it cites.
H. Inaguma, J. Cho, M. K. Baskar, T. Kawahara, and S. Watanabe, “Transfer learning of language-independent end-to-end asr with language model fusion,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6096–6100
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “Specaugment: A simple augmentation method for automatic speech recognition,” in INTERSPEECH , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
N. Rossenbach, A. Zeyer, R. Schlüter, and H. Ney, “Generating synthetic audio data for attention-based speech recognition systems,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7069–7073
2020
Later among the works it cites.
G. Wang, A. Rosenberg, Z. Chen, Y. Zhang, B. Ramabhadran, Y. Wu, and P. Moreno, “Improving speech recognition using consistent predictions on synthesized speech,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7029–7033
2020
Later among the works it cites.