Fetching the paper…
Reading the bibliography…
Data augmentation is one of the most effective ways to make end-to-end automatic speech recognition (ASR) perform close to the conventional hybrid approach, especially when dealing with low-resource tasks.
D. Griffin, D. Deadrick, and Jae Lim, “Speech synthesis from short-time Fourier transform magnitude and its application to speech processing,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , vol. 9, San Diego, CA, USA, 1984, pp. 61–64
1984
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in International conference on Machine learning - ICML , Pittsburgh, Pennsylvania, 2006, pp. 369–376
2006
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” in International Conference on Learning Representations (ICLR) , Dec 2014
2014
Earlier work this paper cites.
T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, “Audio augmentation for speech recognition.” in Interspeech , 2015, pp. 3586–3589
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , South Brisbane, Queensland, Australia, Apr. 2015, pp. 5206–5210
2015
Earlier work this paper cites.
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in Advances in neural information processing systems , 2015, pp. 577–585
2015
Earlier work this paper cites.
T. Drugman, J. Pylkkönen, and R. Kneser, “Active and semi-supervised learning in asr: Benefits on the acoustic and language models,” in Interspeech , Sep. 2016, pp. 2318–2322
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit et al. , “Attention is all you need,” in Advances in Neural Information Processing Systems 30 , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster et al. , “Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 4779–4783
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Y. Wang, D. Stanton, Y. Zhang, R. Skerry-Ryan et al. , “Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,” in International Conference on Machine Learning (ICML) , vol. 80, Jul. 2018, pp. 5167–5176
2018
Cited alongside, same era.
S. Watanabe, T. Hori, S. Karita, T. Hayashi et al. , “ESPnet: End-to-End Speech Processing Toolkit,” in Interspeech , Sep. 2018, pp. 2207–2211
2018
Cited alongside, same era.
T. Kudo and J. Richardson, “SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing,” in Conference on Empirical Methods in Natural Language Processing: System Demonstrations , Brussels, Belgium, 2018, pp. 66–71
2018
Cited alongside, same era.
C. Lüscher, E. Beck, K. Irie, M. Kitza et al. , “RWTH ASR Systems for LibriSpeech: Hybrid vs Attention,” in Interspeech , Sep. 2019, pp. 231–235
2019
Cited alongside, same era.
W.-N. Hsu, Y. Zhang, R. J. Weiss, H. Zen et al. , “Hierarchical generative modeling for controllable speech synthesis,” in International Conference on Learning Representations (ICLR) , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
P. Govalkar, J. Fischer, F. Zalkow, and C. Dittmar, “A comparison of recent neural vocoders for speech signal reconstruction,” in Proc. 10th ISCA Speech Synthesis Workshop , 2019, pp. 7–12
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Wiesner, A. Renduchintala, S. Watanabe, C. Liu et al. , “Pretraining by Backtranslation for End-to-End ASR in Low-Resource Settings,” in Interspeech , 2019, pp. 4375–4379
2019
Cited alongside, same era.
2019
Cited alongside, same era.
A. Rosenberg, Y. Zhang, B. Ramabhadran, Y. Jia et al. , “Speech Recognition with Augmented Synthesized Speech,” in IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , SG, Singapore, Dec. 2019, pp. 996–1002
2019
Cited alongside, same era.
A. Tjandra, S. Sakti, and S. Nakamura, “End-to-end Feedback Loss in Speech Chain Framework via Straight-through Estimator,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Brighton, United Kingdom, May 2019, pp. 6281–6285
2019
Cited alongside, same era.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu et al. , “Specaugment: A simple data augmentation method for automatic speech recognition,” Interspeech , Sep 2019
2019
Cited alongside, same era.
J.-M. Valin and J. Skoglund, “Lpcnet: Improving neural speech synthesis through linear prediction,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 5891–5895
2019
Cited alongside, same era.
S. Karita, X. Wang, S. Watanabe, T. Yoshimura et al. , “A Comparative Study on Transformer vs RNN in Speech Applications,” in IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , SG, Singapore, Dec. 2019, pp. 449–456
2019
Cited alongside, same era.
R. Korostik, A. Chirkovskiy, A. Svischev, I. Kalinovskiy, and A. Talanov, “The stc text-to-speech system for blizzard challenge 2019,” in Proceedings of Blizzard Challenge 2019 , 2019
2019
Later among the works it cites.
2020
Closest in time.
I. Medennikov, M. Korenevsky, T. Prisyach, Y. Khokhlov et al. , “The stc system for the chime-6 challenge,” in CHiME 2020 Workshop on Speech Processing in Everyday Environments , 2020
2020
Closest in time.
N. Rossenbach, A. Zeyer, R. Schluter, and H. Ney, “Generating Synthetic Audio Data for Attention-Based Speech Recognition Systems,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Barcelona, Spain, May 2020, pp. 7069–7073
2020
Closest in time.
G. Sun, Y. Zhang, R. J. Weiss, Y. Cao et al. , “Generating diverse and natural text-to-speech samples using a quantized fine-grained vae and autoregressive prosody prior,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 6699–6703
2020
Closest in time.
E. Battenberg, R. Skerry-Ryan, S. Mariooryad, D. Stanton et al. , “Location-relative attention mechanisms for robust long-form speech synthesis,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 6194–6198
2020
Closest in time.