Fetching the paper…
Reading the bibliography…
Speaker anonymization aims to conceal a speaker's identity while preserving content information in speech.
A. S. Householder, “Unitary triangularization of a nonsymmetric matrix,” Journal of the ACM (JACM) , vol. 5, no. 4, pp. 339–342, 1958
1958
Earlier work this paper cites.
S. E. McAdams, Spectral fusion, spectral parsing and the formation of auditory images . Stanford university, 1984
1984
Earlier work this paper cites.
D. A. Reynolds, “Speaker identification and verification using gaussian mixture speaker models,” Speech communication , vol. 17, no. 1-2, pp. 91–108, 1995
1995
Earlier work this paper cites.
A. A. Dibazar, S. Narayanan, and T. W. Berger, “Feature analysis for automatic detection of pathological speech,” in Proceedings of the second joint 24th annual conference and the annual fall meeting of the biomedical engineering society][engineering in medicine and biology , vol. 1. IEEE, 2002, pp. 182–183
2002
Earlier work this paper cites.
K. Kasi and S. A. Zahorian, “Yet another algorithm for pitch tracking,” in Proc. ICASSP , vol. 1, 2002, pp. I–361
2002
Earlier work this paper cites.
S. Ioffe, “Probabilistic linear discriminant analysis,” in Computer Vision–ECCV 2006: 9th European Conference on Computer Vision, Graz, Austria, May 7-13, 2006, Proceedings, Part IV 9 . Springer Berlin Heidelberg, 2006, pp. 531–542
2006
Earlier work this paper cites.
Q. Jin, A. R. Toth, A. W. Black, and T. Schultz, “Is voice transformation a threat to speaker identification?” in Proc. ICASSP , 2008, pp. 4845–4848
2008
Earlier work this paper cites.
L. Van der Maaten and G. Hinton, “Visualizing data using t-sne.” Journal of machine learning research , vol. 9, no. 11, 2008
2008
Earlier work this paper cites.
Q. Jin, A. R. Toth, T. Schultz, and A. W. Black, “Voice convergin: Speaker de-identification by voice transformation,” in Proc. ICASSP , 2009, pp. 3909–3912
2009
Earlier work this paper cites.
D. Garcia-Romero and C. Y. Espy-Wilson, “Analysis of i-vector length normalization in speaker recognition systems,” in Proc. Interspeech , 2011, pp. 249–254
2011
Earlier work this paper cites.
B. Schuller, S. Steidl, A. Batliner, A. Vinciarelli, K. Scherer, F. Ringeval, M. Chetouani, F. Weninger, F. Eyben, E. Marchi et al. , “The interspeech 2013 computational paralinguistics challenge: Social signals, conflict, emotion, autism,” in Proceedings INTERSPEECH 2013, 14th Annual Conference of the International Speech Communication Association, Lyon, France , 2013
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
V. Peddinti, D. Povey, and S. Khudanpur, “A time delay neural network architecture for efficient modeling of long temporal contexts,” in Proc. Interspeech , 2015, pp. 3214–3218
2015
Earlier work this paper cites.
C. Magarinos, P. Lopez-Otero, L. Docio-Fernandez, E. Rodriguez-Banga, D. Erro, and C. Garcia-Mateo, “Reversible speaker de-identification using pre-trained transformation functions,” Computer Speech & Language , vol. 46, pp. 36–52, 2017
2017
Earlier work this paper cites.
H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, “Aishell-1: An open-source Mandarin speech corpus and a speech recognition baseline,” in 2017 20th Conference of the Oriental Chapter of the International Coordinating Committee on Speech Databases and Speech I/O Systems and Assessment (O-COCOSDA) . IEEE, 2017, pp. 1–5
2017
Earlier work this paper cites.
L. N. Smith, “Cyclical learning rates for training neural networks,” in 2017 IEEE winter conference on applications of computer vision (WACV) . IEEE, 2017, pp. 464–472
2017
Earlier work this paper cites.
J. Qian, H. Du, J. Hou, L. Chen, T. Jung, and X.-Y. Li, “Hidebehind: Enjoy voice input with voiceprint unclonability and anonymity,” in Proceedings of the 16th ACM Conference on Embedded Networked Sensor Systems , 2018, pp. 82–94
2018
Earlier work this paper cites.
D. Povey, G. Cheng, Y. Wang, K. Li, H. Xu, M. Yarmohammadi, and S. Khudanpur, “Semi-orthogonal low-rank matrix factorization for deep neural networks.” in Proc. Interspeech , 2018, pp. 3743–3747
2018
Earlier work this paper cites.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust dnn embeddings for speaker recognition,” in Proc. ICASSP . IEEE, 2018, pp. 5329–5333
2018
Earlier work this paper cites.
A. Kessy, A. Lewin, and K. Strimmer, “Optimal whitening and decorrelation,” The American Statistician , vol. 72, no. 4, pp. 309–314, 2018
2018
Earlier work this paper cites.
J. S. Chung, A. Nagrani, and A. Zisserman, “Voxceleb2: Deep speaker recognition,” in Proc. Interspeech , 2018, pp. 1086–1090
2018
Earlier work this paper cites.
A. Ali, S. Shon, Y. Samih, H. Mubarak, A. Abdelali, J. Glass, S. Renals, and K. Choukri, “The mgb-5 challenge: Recognition and dialect identification of dialectal arabic speech,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2019, pp. 1026–1033
2019
Earlier work this paper cites.
F. Fang, X. Wang, J. Yamagishi, I. Echizen, M. Todisco, N. Evans, and J.-F. Bonastre, “Speaker anonymization using x-vector and neural waveform models,” Proc. 10th ISCA Speech Synthesis Workshop , pp. 155–160, 9 2019
2019
Earlier work this paper cites.
X. Wang, S. Takaki, and J. Yamagishi, “Neural source-filter-based waveform model for statistical parametric speech synthesis,” in Proc. ICASSP , 2019, pp. 5916–5920
2019
Earlier work this paper cites.
J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “Arcface: Additive angular margin loss for deep face recognition,” in Proc. CVPR , 2019, pp. 4690–4699
2019
Cited alongside, same era.
X. Xiang, S. Wang, H. Huang, Y. Qian, and K. Yu, “Margin matters: Towards more discriminative deep neural network embeddings for speaker recognition,” in Proc. APSIPA ASC . IEEE, 2019, pp. 1652–1656
2019
Cited alongside, same era.
2019
Cited alongside, same era.
C. Veaux, J. Yamagishi, and K. MacDonald, “CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92),” 2019. [Online]. Available: https://datashare.is.ed.ac.uk/handle/10283/3443
2019
Cited alongside, same era.
N. Tomashenko, X. Wang, E. Vincent, J. Patino, B. M. L. Srivastava, P.-G. Noé, A. Nautsch, N. Evans, J. Yamagishi, B. O’Brien et al. , “The VoicePrivacy 2020 challenge: Results and findings,” Computer Speech & Language , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
L. Tavi, T. Kinnunen, and R. G. Hautamäki, “Improving speaker de-identification with functional data analysis of f0 trajectories,” Speech Communication , vol. 140, pp. 1–10, 2022
2022
Later among the works it cites.
C. O. Mawalim, S. Okada, and M. Unoki. System description: Speaker anonymization by pitch shifting based on time-scale modification (pv-tsm). [Online]. Available: https://www.voiceprivacychallenge.org/results-2022/docs/1___T32.pdf
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al. , “Pytorch: An imperative style, high-performance deep learning library,” Advances in neural information processing systems , vol. 32, 2019
2019
Cited alongside, same era.
V. Vestman, T. Kinnunen, R. G. Hautamäki, and M. Sahidullah, “Voice mimicry attacks assisted by automatic speaker verification,” Computer Speech & Language , vol. 59, pp. 36–54, 2020
2020
Cited alongside, same era.
R. K. Das, X. Tian, T. Kinnunen, and H. Li, “The Attacker’s Perspective on Automatic Speaker Verification: An Overview,” in Proc. Interspeech 2020 , 2020, pp. 4213–4217. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2020-1052
2020
Cited alongside, same era.
N. Tomashenko, B. M. L. Srivastava, X. Wang, E. Vincent, A. Nautsch, J. Yamagishi, N. Evans, J. Patino, J.-F. Bonastre, P.-G. Noé, and M. Todisco, “Introducing the VoicePrivacy Initiative,” in Proc. Interspeech , 2020, pp. 1693–1697
2020
Cited alongside, same era.
P. Gupta, G. P. Prajapati, S. Singh, M. R. Kamble, and H. A. Patil, “Design of voice privacy system using linear prediction,” in 2020 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) . IEEE, 2020, pp. 543–549
2020
Cited alongside, same era.
S. P. Dubagunta, R. Van Son, and M. M. Doss, “Adjustable deterministic pseudonymisation of speech: Idiap-nki’s submission to VoicePrivacy 2020 challenge,” URL: https://www. voiceprivacychallenge. org/docs/Idiap-NKI. pdf , 2020
2020
Cited alongside, same era.
B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification,” in Proc. Interspeech , 2020, pp. 3830–3834
2020
Cited alongside, same era.
B. M. L. Srivastava, N. A. Tomashenko, X. Wang, E. Vincent, J. Yamagishi, M. Maouche, A. Bellet, and M. Tommasi, “Design choices for x-vector based speaker anonymization,” in Proc. Interspeech , 2020, pp. 1713–1717
2020
Cited alongside, same era.
B. M. L. Srivastava, M. Maouche, M. Sahidullah, E. Vincent, A. Bellet, M. Tommasi, N. Tomashenko, X. Wang, and J. Yamagishi, “Privacy and utility of x-vector based speaker anonymization,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 2383–2395, 2022
2022
Later among the works it cites.
S. Meyer, P. Tilli, P. Denisov, F. Lux, J. Koch, and N. T. Vu, “Anonymizing speech with generative adversarial networks to preserve speaker privacy,” in Proc. Spoken Language Technology Workshop (SLT) , 2022, p. To be published
2022
Later among the works it cites.
J. Yao, Q. Wang, L. Zhang, P. Guo, Y. Liang, and L. Xie, “NWPU-ASLP System for the VoicePrivacy 2022 Challenge,” in Proc. 2nd Symposium on Security and Privacy in Speech Communication , 2022
2022
Later among the works it cites.
X. Miao, X. Wang, E. Cooper, J. Yamagishi, and N. Tomashenko, “Language-Independent Speaker Anonymization Approach Using Self-Supervised Pre-Trained Models,” in Proc. The Speaker and Language Recognition Workshop (Odyssey 2022) , 2022, pp. 279–286
2022
Later among the works it cites.
——, “Analyzing Language-Independent Speaker Anonymization Framework under Unseen Conditions,” in Proc. Interspeech 2022 , 2022, pp. 4426–4430
2022
Later among the works it cites.
X. Chen, G. Li, H. Huang, W. Zhou, S. Li, Y. Cao, and Y. Zhao, “System description for Voice Privacy Challenge 2022 ,” in Proc. 2nd Symposium on Security and Privacy in Speech Communication , 2022
2022
Later among the works it cites.
U. E. Gaznepoglu, A. Leschanowsky, and N. Peters, “VoicePrivacy 2022 system description: speaker anonymization with feature-matched f0 trajectories,” in Proc. 2nd Symposium on Security and Privacy in Speech Communication , 2022
2022
Later among the works it cites.
C. O. Mawalim, S. Okada, and M. Unoki, “Speaker anonymization by pitch shifting based on time-scale modification,” in Proc. 2nd Symposium on Security and Privacy in Speech Communication , 2022, pp. 35–42
2022
Later among the works it cites.
R. Khamsehashari, Y. Sinha, J. Hintz, S. Ghosh, T. Polzehl, C. Franzreb, S. Stober, and I. Siegert, “Voice Privacy - leveraging multi-scale blocks with ECAPA-TDNN SE-Res2NeXt extension for speaker anonymization,” in Proc. 2nd Symposium on Security and Privacy in Speech Communication , 2022, pp. 43–48
2022
Later among the works it cites.
P.-G. Noé, A. Nautsch, N. Evans, J. Patino, J.-F. Bonastre, N. Tomashenko, and D. Matrouf, “Towards a unified assessment framework of speech pseudonymisation,” Computer Speech & Language , vol. 72, p. 101299, 2022
2022
Later among the works it cites.
C. O. Mawalim, K. Galajit, J. Karnjana, S. Kidani, and M. Unoki, “Speaker anonymization by modifying fundamental frequency and x-vector singular value,” Computer Speech & Language , vol. 73, p. 101326, 2022
2022
Later among the works it cites.
C. Pierre, A. Larcher, and D. Jouvet, “Are disentangled representations all you need to build speaker anonymization systems?” in Proc. Interspeech 2022 , 2022, pp. 2793–2797
2022
Later among the works it cites.
J. M. Perero-Codosero, F. M. Espinoza-Cuadros, and L. A. Hernández-Gómez, “X-vector anonymization using autoencoders and adversarial training for preserving speech privacy,” Computer Speech & Language , vol. 74, p. 101351, 2022
2022
Later among the works it cites.
H. Turner, G. Lovisotto, and I. Martinovic, “Generating identities with mixture models for speaker anonymization,” Computer Speech & Language , vol. 72, p. 101318, 2022
2022
Later among the works it cites.
B. van Niekerk, M.-A. Carbonneau, J. Zaïdi, M. Baas, H. Seuté, and H. Kamper, “A comparison of discrete and soft speech units for improved voice conversion,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 6562–6566
2022
Later among the works it cites.
Q. Wang, K. A. Lee, and T. Liu, “Scoring of Large-Margin Embeddings for Speaker Verification: Cosine or PLDA?” in Proc. Interspeech 2022 , 2022, pp. 600–604
2022
Later among the works it cites.
L. Li, R. Liu, J. Kang, Y. Fan, H. Cui, Y. Cai, R. Vipperla, T. F. Zheng, and D. Wang, “CN-Celeb: multi-genre speaker recognition,” Speech Communication , 2022
2022
Later among the works it cites.
E. Cooper, W.-C. Huang, T. Toda, and J. Yamagishi, “Generalization ability of MOS prediction networks,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 8442–8446
2022
Later among the works it cites.
A. S. Shamsabadi, B. M. L. Srivastava, A. Bellet, N. Vauquier, E. Vincent, M. Maouche, M. Tommasi, and N. Papernot, “Differentially private speaker anonymization,” Proceedings on Privacy Enhancing Technologies , vol. 2023, no. 1, Jan. 2023. [Online]. Available: https://hal.inria.fr/hal-03588932
2023
Closest in time.
P.-G. Noé, X. Miao, X. Wang, J. Yamagishi, J.-F. Bonastre, and D. Matrouf, “Hiding speaker’s sex in speech using zero-evidence speaker representation in an analysis/synthesis pipeline,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Closest in time.