Fetching the paper…
Reading the bibliography…
Speaker anonymization aims to protect the privacy of speakers while preserving spoken linguistic information from speech.
Spectral fusion, spectral parsing and the formation of auditory images
Stephen Edward McAdams, · 1984
Earlier work this paper cites.
“Yet another algorithm for pitch tracking,”
Kavita Kasi and Stephen A Zahorian, · 2002
Earlier work this paper cites.
“Is voice transformation a threat to speaker identification?,”
Qin Jin, Arthur R Toth, Alan W Black, and Tanja Schultz, · 2008
Earlier work this paper cites.
“Voice convergin: Speaker de-identification by voice transformation,”
Qin Jin, Arthur R Toth, Tanja Schultz, and Alan W Black, · 2009
Earlier work this paper cites.
“A time delay neural network architecture for efficient modeling of long temporal contexts,”
Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Librispeech: an ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Privacy-preserving sound to degrade automatic speaker verification performance,”
Kei Hashimoto, Junichi Yamagishi, and Isao Echizen, · 2016
Earlier work this paper cites.
“End-to-end text-dependent speaker verification,”
Georg Heigold, Ignacio Moreno, Samy Bengio, and Noam Shazeer, · 2016
Earlier work this paper cites.
“Voicemask: Anonymize and sanitize voice input on mobile devices,”
Jianwei Qian, Haohua Du, Jiahui Hou, Linlin Chen, Taeho Jung, Xiangyang Li, Yu Wang, and Yanbo Deng, · 2017
Earlier work this paper cites.
“Reversible speaker de-identification using pre-trained transformation functions,”
Carmen Magarinos, Paula Lopez-Otero, Laura Docio-Fernandez, Eduardo Rodriguez-Banga, Daniel Erro, and Carmen Garcia-Mateo, · 2017
Earlier work this paper cites.
“Voxceleb: A large-scale speaker identification dataset,”
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman, · 2017
Earlier work this paper cites.
“Aishell-1: An open-source Mandarin speech corpus and a speech recognition baseline,”
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng, · 2017
Earlier work this paper cites.
“Hidebehind: Enjoy voice input with voiceprint unclonability and anonymity,”
Jianwei Qian, Haohua Du, Jiahui Hou, Linlin Chen, Taeho Jung, and Xiang-Yang Li, · 2018
Earlier work this paper cites.
“X-vectors: Robust dnn embeddings for speaker recognition,”
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur, · 2018
Earlier work this paper cites.
“Semi-orthogonal low-rank matrix factorization for deep neural networks.,”
Daniel Povey, Gaofeng Cheng, Yiming Wang, Ke Li, Hainan Xu, Mahsa Yarmohammadi, and Sanjeev Khudanpur, · 2018
Earlier work this paper cites.
“Voxceleb2: Deep speaker recognition,”
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman, · 2018
Earlier work this paper cites.
“Exploring the encoding layer and loss function in end-to-end speaker and language recognition system,”
Weicheng Cai, Jinkun Chen, and Ming Li, · 2018
Earlier work this paper cites.
“Speaker anonymization using x-vector and neural waveform models,”
Fuming Fang, Xin Wang, Junichi Yamagishi, Isao Echizen, Massimiliano Todisco, Nicholas Evans, and Jean-Francois Bonastre, · 2019
Cited alongside, same era.
“Privacy-preserving adversarial representation learning in ASR: Reality or illusion?,”
Brij Mohan Lal Srivastava, Aurélien Bellet, Marc Tommasi, and Emmanuel Vincent, · 2019
Cited alongside, same era.
“Neural source-filter-based waveform model for statistical parametric speech synthesis,”
Xin Wang, Shinji Takaki, and Junichi Yamagishi, · 2019
Cited alongside, same era.
“LibriTTS: A corpus derived from LibriSpeech for text-to-speech,”
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu, · 2019
Cited alongside, same era.
“Arcface: Additive angular margin loss for deep face recognition,”
Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou, · 2019
Cited alongside, same era.
“Defending your voice: Adversarial attack on voice conversion,”
Chien-yu Huang, Yist Y Lin, Hung-yi Lee, and Lin-shan Lee, · 2021
Later among the works it cites.
“Generative spoken language modeling from raw audio,”
Kushal Lakhotia, Evgeny Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Adelrahman Mohamed, et al., · 2021
Later among the works it cites.
“Speech resynthesis from discrete disentangled self-supervised representations,”
Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux, · 2021
Later among the works it cites.
“Any-to-one sequence-to-sequence voice conversion using self-supervised discrete speech representations,”
Wen-Chin Huang, Yi-Chiao Wu, and Tomoki Hayashi, · 2021
Later among the works it cites.
“Neural analysis and synthesis: Reconstructing speech from self-supervised representations,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Margin matters: Towards more discriminative deep neural network embeddings for speaker recognition,”
Xu Xiang, Shuai Wang, Houjun Huang, Yanmin Qian, and Kai Yu, · 2019
Cited alongside, same era.
“CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92),” 2019
Christophe Veaux, Junichi Yamagishi, and Kirsten MacDonald, · 2019
Cited alongside, same era.
“Probing the information encoded in x-vectors,”
Desh Raj, David Snyder, Daniel Povey, and Sanjeev Khudanpur, · 2019
Cited alongside, same era.
“Disentangling style factors from speaker representations.,”
Jennifer Williams and Simon King, · 2019
Cited alongside, same era.
“Voice mimicry attacks assisted by automatic speaker verification,”
Ville Vestman, Tomi Kinnunen, Rosa González Hautamäki, and Md Sahidullah, · 2020
Cited alongside, same era.
“Introducing the VoicePrivacy Initiative,”
N. Tomashenko, Brij Mohan Lal Srivastava, Xin Wang, Emmanuel Vincent, Andreas Nautsch, Junichi Yamagishi, Nicholas Evans, Jose Patino, Jean-François Bonastre, Paul-Gauthier Noé, and Massimiliano Todisco, · 2020
Cited alongside, same era.
“ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification,”
Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck, · 2020
Cited alongside, same era.
Hyeong-Seok Choi, Juheon Lee, Wansoo Kim, Jie Lee, Hoon Heo, and Kyogu Lee, · 2021
Later among the works it cites.
“A comparison of discrete and soft speech units for improved voice conversion,”
Benjamin van Niekerk, Marc-André Carbonneau, Julian Zaïdi, Mathew Baas, Hugo Seuté, and Herman Kamper, · 2021
Later among the works it cites.
“AISHELL-3: A Multi-Speaker Mandarin TTS Corpus,”
Yao Shi, Hui Bu, Xin Xu, Shaoji Zhang, and Ming Li, · 2021
Later among the works it cites.
“Speaker Anonymisation Using the McAdams Coefficient,”
Jose Patino, Natalia Tomashenko, Massimiliano Todisco, Andreas Nautsch, and Nicholas Evans, · 2021
Later among the works it cites.
“HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Later among the works it cites.
“Layer-wise Analysis of a Self-supervised Speech Representation Model,”
Ankita Pasad, Ju-Chieh Chou, and Karen Livescu, · 2021
Later among the works it cites.
“SUPERB: Speech Processing Universal PERformance Benchmark,”
Shu wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y. Lin, Andy T. Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, Tzu-Hsien Huang, Wei-Cheng Tseng, Ko tik Lee, Da-Rong Liu, Zili Huang, Shuyan Dong, Shang-Wen Li, Shinji Watanabe, Abdelrahman Mohamed, and Hung yi Lee, · 2021
Later among the works it cites.
“D-MONA: A dilated mixed-order non-local attention network for speaker and language recognition,”
Xiaoxiao Miao, Ian McLoughlin, Wenchao Wang, and Pengyuan Zhang, · 2021
Later among the works it cites.
“SpeechBrain: A general-purpose speech toolkit,” 2021,
Mirco Ravanelli, Titouan Parcollet, Peter Plantinga, Aku Rouhe, Samuele Cornell, Loren Lugosch, Cem Subakan, Nauman Dawalatabad, Abdelwahab Heba, Jianyuan Zhong, Ju-Chieh Chou, Sung-Lin Yeh, Szu-Wei Fu, Chien-Feng Liao, Elena Rastorgueva, François Grondin, William Aris, Hwidong Na, Yan Gao, Renato De Mori, and Yoshua Bengio, · 2021
Later among the works it cites.
“The VoicePrivacy 2020 challenge: Results and findings,”
Natalia Tomashenko, Xin Wang, Emmanuel Vincent, Jose Patino, Brij Mohan Lal Srivastava, Paul-Gauthier Noé, Andreas Nautsch, Nicholas Evans, Junichi Yamagishi, Benjamin O’Brien, et al., · 2022
Closest in time.
“Large-scale self-supervised speech representation learning for automatic speaker verification,”
Zhengyang Chen, Sanyuan Chen, Yu Wu, Yao Qian, Chengyi Wang, Shujie Liu, Yanmin Qian, and Michael Zeng, · 2022
Closest in time.
“Generalization ability of MOS prediction networks,”
Erica Cooper, Wen-Chin Huang, Tomoki Toda, and Junichi Yamagishi, · 2022
Closest in time.
“CN-Celeb: multi-genre speaker recognition,”
Lantian Li, Ruiqi Liu, Jiawen Kang, Yue Fan, Hao Cui, Yunqi Cai, Ravichander Vipperla, Thomas Fang Zheng, and Dong Wang, · 2022
Closest in time.