Fetching the paper…
Reading the bibliography…
For new participants - Executive summary: (1) The task is to develop a voice anonymization system for speech data which conceals the speaker's voice identity while protecting linguistic content, paralinguistic attributes, intelligibility and naturalness.
S. McAdams, “Spectral fusion, spectral parsing and the formation of the auditory image,” Ph.D. dissertation, Stanford University, 1984
1984
Earlier work this paper cites.
A. Martin, G. Doddington, T. Kamm, M. Ordowski, and M. Przybocki, “The DET curve in assessment of detection task performance,” National Inst of Standards and Technology Gaithersburg MD, Tech. Rep., 1997
1997
Earlier work this paper cites.
D. Hirst, “A Praat plugin for momel and intsint with improved algorithms for modelling and coding intonation. icphs xvi, saabrücken,” 2007
2007
Earlier work this paper cites.
S. Ghorshi, S. Vaseghi, and Q. Yan, “Cross-entropic comparison of formants of British, Australian and American English accents,” Speech Communication , vol. 50, no. 7, pp. 564–579, 2008
2008
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel et al. , “The Kaldi speech recognition toolkit,” 2011
2011
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “LibriSpeech: an ASR corpus based on public domain audio books,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015, pp. 5206–5210
2015
Earlier work this paper cites.
V. Peddinti, D. Povey, and S. Khudanpur, “A time delay neural network architecture for efficient modeling of long temporal contexts,” in Interspeech , 2015, pp. 3214–3218
2015
Earlier work this paper cites.
A. Larcher, K. A. Lee, and S. Meignier, “An extensible speaker identification sidekit in python,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2016, pp. 5095–5099
2016
Earlier work this paper cites.
A. Nagrani, J. S. Chung, and A. Zisserman, “VoxCeleb: a large-scale speaker identification dataset,” in Interspeech , 2017, pp. 2616–2620
2017
Earlier work this paper cites.
J. Qian, F. Han, J. Hou, C. Zhang, Y. Wang, and X.-Y. Li, “Towards privacy-preserving speech data publishing,” in 2018 IEEE Conference on Computer Communications (INFOCOM) , 2018, pp. 1079–1087
2018
Earlier work this paper cites.
J. S. Chung, A. Nagrani, and A. Zisserman, “VoxCeleb2: Deep speaker recognition,” in Interspeech , 2018, pp. 1086–1090
2018
Earlier work this paper cites.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust DNN embeddings for speaker recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 5329–5333
2018
Earlier work this paper cites.
D. Povey, G. Cheng, Y. Wang, K. Li, H. Xu, M. Yarmohammadi et al. , “Semi-orthogonal low-rank matrix factorization for deep neural networks.” in Interspeech , 2018, pp. 3743–3747
2018
Earlier work this paper cites.
A. Nautsch, C. Jasserand, E. Kindt, M. Todisco, I. Trancoso, and N. Evans, “The GDPR & speech data: Reflections of legal and technology communities, first steps towards a common understanding,” in Interspeech , 2019, pp. 3695–3699
2019
Earlier work this paper cites.
A. Nautsch, A. Jimenez, A. Treiber, J. Kolberg, C. Jasserand, E. Kindt, H. Delgado et al. , “Preserving privacy in speaker and speech characterisation,” Computer Speech and Language , vol. 58, pp. 441–480, 2019
2019
Cited alongside, same era.
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu, “LibriTTS: A corpus derived from LibriSpeech for text-to-speech,” in Interspeech , 2019, pp. 1526–1530
2019
Cited alongside, same era.
C. Veaux, J. Yamagishi, and K. MacDonald, “CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92),” https://datashare.is.ed.ac.uk/handle/10283/3443 , 2019
2019
Cited alongside, same era.
F. Fang, X. Wang, J. Yamagishi, I. Echizen, M. Todisco, N. Evans, and J.-F. Bonastre, “Speaker anonymization using x-vector and neural waveform models,” in Speech Synthesis Workshop , 2019, pp. 155–160
2019
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in Neural Information Processing Systems , vol. 33, pp. 12 449–12 460, 2020
2020
Later among the works it cites.
——, “Supplementary material to the paper. The VoicePrivacy 2020 Challenge: Results and findings,” https://hal.archives-ouvertes.fr/hal-03335126 , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
B. M. L. Srivastava, M. Maouche, M. Sahidullah, E. Vincent, A. Bellet, M. Tommasi, N. Tomashenko, X. Wang, and J. Yamagishi, “Privacy and utility of x-vector based speaker anonymization,” 2021. [Online]. Available: https://hal.inria.fr/hal-03197376/file/design_choices_informed.pdf
2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Wang and J. Yamagishi, “Neural harmonic-plus-noise waveform model with trainable maximum voice frequency for text-to-speech synthesis,” in Speech Synthesis Workshop , 2019, pp. 1–6
2019
Cited alongside, same era.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al. , “Pytorch: An imperative style, high-performance deep learning library,” in Advances in Neural Information Processing Systems , H. Wallach et al. , Eds., vol. 32. Curran Associates, Inc., 2019. [Online]. Available: https://proceedings.neurips.cc/paper/2019/file/bdbca288fee7f92f2bfa9f7012727740-Paper.pdf
2019
Cited alongside, same era.
N. Tomashenko, B. M. L. Srivastava, X. Wang, E. Vincent, A. Nautsch, J. Yamagishi, N. Evans, J. Patino, J.-F. Bonastre, P.-G. Noé, and M. Todisco, “Introducing the VoicePrivacy Initiative,” in Proc. Interspeech 2020 , 2020, pp. 1693–1697. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2020-1333
2020
Cited alongside, same era.
B. M. L. Srivastava, N. Vauquier, M. Sahidullah, A. Bellet, M. Tommasi, and E. Vincent, “Evaluating voice conversion-based privacy protection against informed attackers,” in 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 2802–2806
2020
Cited alongside, same era.
N. Tomashenko, B. M. L. Srivastava, X. Wang, E. Vincent, A. Nautsch, J. Yamagishi, N. Evans, J. Patino, J.-F. Bonastre, P.-G. Noé, and M. Todisco, “The VoicePrivacy 2020 Challenge evaluation plan,” https://www.voiceprivacychallenge.org/docs/VoicePrivacy_2020_Eval_Plan_v1_3.pdf , 2020
2020
Cited alongside, same era.
N. Tomashenko, B. M. L. Srivastava, X. Wang, E. Vincent, A. Nautsch, J. Yamagishi, N. Evans et al. , “Post-evaluation analysis for the VoicePrivacy 2020 challenge: Using anonymized speech data to train attack models and ASR,” https://www.voiceprivacychallenge.org/docs/VoicePrivacy2020_post_evaluation.pdf , 2020
2020
Cited alongside, same era.
P.-G. Noé, J.-F. Bonastre, D. Matrouf, N. Tomashenko, A. Nautsch, and N. Evans, “Speech pseudonymisation assessment using voice similarity matrices,” in Interspeech , 2020, pp. 1718–1722
2020
Cited alongside, same era.
J. Kong, J. Kim, and J. Bae, “Hifi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” Advances in Neural Information Processing Systems , vol. 33, pp. 17 022–17 033, 2020
2020
Cited alongside, same era.
Later among the works it cites.
M. Maouche, B. M. L. Srivastava, N. Vauquier, A. Bellet, M. Tommasi, and E. Vincent, “Enhancing speech privacy with slicing,” 2021. [Online]. Available: https://hal.inria.fr/hal-03369137
2021
Later among the works it cites.
J. Patino, N. Tomashenko, M. Todisco, A. Nautsch, and N. Evans, “Speaker anonymisation using the McAdams coefficient,” in Interspeech , 2021, pp. 1099–1103
2021
Later among the works it cites.
J. Deng, J. Guo, J. Yang, N. Xue, I. Cotsia, and S. Zafeiriou, “Arcface: Additive angular margin loss for deep face recognition.” IEEE transactions on pattern analysis and machine intelligence , vol. PP, 2021
2021
Later among the works it cites.
C. Wang, M. Riviere, A. Lee, A. Wu, C. Talnikar, D. Haziza, M. Williamson, J. Pino, and E. Dupoux, “VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . Online: Association for Computational Linguistics, Aug. 2021, pp. 993–1003. [Online]. Available: https://aclanthology.org/2021.acl-long.80
2021
Later among the works it cites.
2022
Closest in time.
2022
Closest in time.
P.-G. Noé, A. Nautsch, N. Evans, J. Patino, J.-F. Bonastre, N. Tomashenko, and D. Matrouf, “Towards a unified assessment framework of speech pseudonymisation,” Computer Speech & Language , vol. 72, p. 101299, 2022
2022
Closest in time.
2022
Closest in time.