Fetching the paper…
Reading the bibliography…
The goal of accent conversion (AC) is to convert the accent of speech into the target accent while preserving the content and speaker identity.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proceedings of the 23rd international conference on Machine learning , 2006, pp. 369–376
2006
Earlier work this paper cites.
C. Xia, C. Xiong, and P. Yu, “Pseudo siamese network for few-shot intent generation,” in Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval , ser. SIGIR ’21. New York, NY, USA: Association for Computing Machinery, 2021, p. 2005–2009. [Online]. Available: https://doi.org/10.1145/3404835.3462995
2009
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Y. Ganin and V. Lempitsky, “Unsupervised domain adaptation by backpropagation,” in International conference on machine learning . PMLR, 2015, pp. 1180–1189
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in International conference on machine learning . pmlr, 2015, pp. 448–456
2015
Earlier work this paper cites.
2017
Earlier work this paper cites.
C. Veaux, J. Yamagishi, K. MacDonald et al. , “Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit,” University of Edinburgh. The Centre for Speech Technology Research (CSTR) , 2017
2017
Earlier work this paper cites.
G. Zhao, S. Sonsaat, J. Levis, E. Chukharev-Hudilainen, and R. Gutierrez-Osuna, “Accent conversion using phonetic posteriorgrams,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 5314–5318
2018
Cited alongside, same era.
H.-Y. Lee, H.-Y. Tseng, J.-B. Huang, M. Singh, and M.-H. Yang, “Diverse image-to-image translation via disentangled representations,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 35–51
2018
Cited alongside, same era.
G. Zhao, S. Sonsaat, A. Silpachai, I. Lucic, E. Chukharev-Hudilainen, J. Levis, and R. Gutierrez-Osuna, “L2-arctic: A non-native english speech corpus,” Proc. Interspeech 2018 , pp. 2783–2787, 2018
2018
Cited alongside, same era.
T. Zenkel, R. Sanabria, F. Metze, and A. Waibel, “Subword and crossword units for ctc acoustic models,” Proc. Interspeech 2018 , pp. 396–400, 2018
2018
Cited alongside, same era.
S. Liu, D. Wang, Y. Cao, L. Sun, X. Wu, S. Kang, Z. Wu, X. Liu, D. Su, D. Yu, and H. Meng, “End-to-end accent conversion without using native utterances,” in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 6289–6293
2020
Later among the works it cites.
J. Gao, C. Xiao, L. Glass, and J. Sun, “Compose: Cross-modal pseudo-siamese network for patient trial matching,” Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , 2020
2020
Later among the works it cites.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented Transformer for Speech Recognition,” in Proc. Interspeech 2020 , 2020, pp. 5036–5040. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2020-3015
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Wang, D. Stanton, Y. Zhang, R.-S. Ryan, E. Battenberg, J. Shor, Y. Xiao, Y. Jia, F. Ren, and R. A. Saurous, “Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,” in International Conference on Machine Learning . PMLR, 2018, pp. 5180–5189
2018
Cited alongside, same era.
G. Zhao, S. Ding, and R. Gutierrez-Osuna, “Foreign Accent Conversion by Synthesizing Speech from Phonetic Posteriorgrams,” in Proc. Interspeech 2019 , 2019, pp. 2843–2847
2019
Cited alongside, same era.
A. Gresse, M. Quillot, R. Dufour, V. Labatut, and J.-F. Bonastre, “Similarity metric based on siamese neural networks for voice casting,” in ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 6585–6589
2019
Cited alongside, same era.
N. Mor, L. Wolf, A. Polyak, and Y. Taigman, “Autoencoder-based music translation,” in International Conference on Learning Representations , 2019. [Online]. Available: https://openreview.net/forum?id=HJGkisCcKm
2019
Cited alongside, same era.
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu, “Libritts: A corpus derived from librispeech for text-to-speech,” Proc. Interspeech 2019 , pp. 1526–1530, 2019
2019
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Later among the works it cites.
S. Ding, G. Zhao, and R. Gutierrez-Osuna, “Accentron: Foreign accent conversion to arbitrary non-native speakers using zero-shot learning,” Comput. Speech Lang. , vol. 72, no. C, mar 2022. [Online]. Available: https://doi.org/10.1016/j.csl.2021.101302
2021
Later among the works it cites.
G. Zhao, S. Ding, and R. Gutierrez-Osuna, “Converting foreign accent speech without a reference,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 2367–2381, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
T. N. Nguyen, N.-Q. Pham, and A. Waibel, “Accent Conversion using Pre-trained Model and Synthesized Data from Voice Conversion,” in Proc. Interspeech 2022 , 2022, pp. 2583–2587
2022
Closest in time.
W. Quamer, A. Das, J. Levis, E. Chukharev-Hudilainen, and R. Gutierrez-Osuna, “Zero-shot foreign accent conversion without a native reference,” in Proc. Interspeech , 2022
2022
Closest in time.
S. Khorram, J. Kim, A. Tripathi, H. Lu, Q. Zhang, and H. Sak, “Contrastive siamese network for semi-supervised speech recognition,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 7207–7211
2022
Closest in time.