Fetching the paper…
Reading the bibliography…
This paper presents a method for selecting appropriate synthetic speech samples from a given large text-to-speech (TTS) dataset as supplementary training data for an automatic speech recognition (ASR) model.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
B. Settles and M. Craven, “An analysis of active learning strategies for sequence labeling tasks,” in Conference on Empirical Methods in Natural Language Processing , 2008, pp. 1070–1079
2008
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Long Beach, CA, USA, 2017
2017
Earlier work this paper cites.
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. van den Oord, S. Dieleman, and K. Kavukcuoglu, “Efficient neural audio synthesis,” in International Conference on Machine Learning (ICML) , 2018, pp. 2410–2419
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Rosenberg, Y. Zhang, B. Ramabhadran, Y. Jia, P. Moreno, Y. Wu, and Z. Wu, “Speech recognition with augmented synthesized speech,” in IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , Sentosa,Singapore, 2019, pp. 996–1002
2019
Earlier work this paper cites.
2019
Cited alongside, same era.
J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “Arcface: Additive angular margin loss for deep face recognition,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , Long Beach, CA, USA, 2019, pp. 4685–4694
2019
Cited alongside, same era.
2019
Cited alongside, same era.
G. Wang, A. Rosenberg, Z. Chen, Y. Zhang, B. Ramabhadran, Y. Wu, and P. Moreno, “Improving speech recognition using consistent predictions on synthesized speech,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Barcelona,Spain, 2020, pp. 7029–7033
C. Wu, Z. Xiu, Y. Shi, O. Kalinli, C. Fuegen, T. Koehler, and Q. He, “Transformer-Based Acoustic Modeling for Streaming Speech Synthesis,” in Proc. Interspeech 2021 , Shanghai,China, 2021, pp. 146–150
2021
Later among the works it cites.
Y. Shi, Y. Wang, C. Wu, C.-F. Yeh, J. Chan, F. Zhang, D. Le, and M. Seltzer, “Emformer: Efficient memory transformer based acoustic model for low latency streaming speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 6783–6787
2021
Later among the works it cites.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3451–3460, 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
S. Ueno, M. Mimura, S. Sakai, and T. Kawahara, “Data augmentation for asr using tts via a discrete representation,” in IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , Cartagena,Colombia, 2021, pp. 68–75
2021
Cited alongside, same era.
Q. He, Z. Xiu, T. Koehler, and J. Wu, “Multi-rate attention architecture for fast streamable text-to-speech spectrum modeling,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Toronto,Canada, 2021, pp. 5689–5693
2021
Cited alongside, same era.
2021
Later among the works it cites.
T. Hu, M. Armandpour, A. Shrivastava, J. R. Chang, H. Koppula, and O. Tuzel, “Synt++: Utilizing imperfect synthetic data to improve speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Singapore,Singapore, 2022, pp. 7682–7686
2022
Later among the works it cites.