Fetching the paper…
Reading the bibliography…
We introduce DISSC, a novel, lightweight method that converts the rhythm, pitch contour and timbre of a recording to a target speaker in a textless manner.
Ju-chieh Chou, Cheng-chieh Yeh, and Hung-yi Lee. 2019 · 1904
Earlier work this paper cites.
The effects of selected factors on the aural identification of speakers
Carl Earl Williams. 1965 · 1965
Earlier work this paper cites.
Spectral voice conversion for text-to-speech synthesis
Alexander Kain and Michael W Macon. 1998 · 1998
Earlier work this paper cites.
A metric for distributions with applications to image databases
Yossi Rubner, Carlo Tomasi, and Leonidas J Guibas. 1998 · 1998
Earlier work this paper cites.
Continuous probabilistic transform for voice conversion
Yannis Stylianou, Olivier Cappé, and Eric Moulines. 1998 · 1998
Earlier work this paper cites.
Yet another algorithm for pitch tracking
Kavita Kasi and Stephen A Zahorian. 2002 · 2002
Earlier work this paper cites.
Voice conversion in high-order eigen space using deep belief nets
Toru Nakashika, Ryoichi Takashima, Tetsuya Takiguchi, and Yasuo Ariki. 2013 · 2013
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
Parallel-data-free voice conversion using cycle-consistent adversarial networks
Takuhiro Kaneko and Hirokazu Kameoka. 2017 · 2017
Earlier work this paper cites.
Montreal Forced Aligner: Trainable Text-Speech Alignment Using Kaldi
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger. 2017 · 2017
Earlier work this paper cites.
Superseded-cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit
Christophe Veaux, Junichi Yamagishi, Kirsten MacDonald, et al. 2017 · 2017
Earlier work this paper cites.
Stargan-vc: Non-parallel many-to-many voice conversion using star generative adversarial networks
Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka, and Nobukatsu Hojo. 2018 · 2018
Earlier work this paper cites.
Autovc: Zero-shot voice style transfer with only autoencoder loss
Kaizhi Qian, Yang Zhang, Shiyu Chang, Xuesong Yang, and Mark Hasegawa-Johnson. 2019 · 2019
Earlier work this paper cites.
ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification
Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck. 2020 · 2020
Earlier work this paper cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020 · 2020
Earlier work this paper cites.
Unsupervised speech decomposition via triple information bottleneck
Kaizhi Qian, Yang Zhang, Shiyu Chang, Mark Hasegawa-Johnson, and David Cox. 2020 · 2020
Cited alongside, same era.
Again-vc: A one-shot voice conversion using activation guidance and adaptive instance normalization
Yen-Hao Chen, Da-Yi Wu, Tsung-Han Wu, and Hung-yi Lee. 2021 · 2021
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021 · 2021
Cited alongside, same era.
Any-to-one sequence-to-sequence voice conversion using self-supervised discrete speech representations
Wen-Chin Huang, Yi-Chiao Wu, and Tomoki Hayashi. 2021 · 2021
Cited alongside, same era.
Textless speech emotion conversion using decomposed and discrete representations
Felix Kreuk, Adam Polyak, Jade Copet, Eugene Kharitonov, Tu-Anh Nguyen, Morgane Rivière, Wei-Ning Hsu, Abdelrahman Mohamed, Emmanuel Dupoux, and Yossi Adi. 2021 · 2021
A unified one-shot prosody and speaker conversion system with self-supervised discrete speech units
Li-Wei Chen, Shinji Watanabe, and Alexander Rudnicky. 2022 · 2022
Closest in time.
Controlvc: Zero-shot voice conversion with time-varying controls on pitch and rhythm
Meiying Chen and Zhiyao Duan. 2022 · 2022
Closest in time.
On the robustness of self-supervised representations for spoken language modeling
Itai Gat, Felix Kreuk, Ann Lee, Jade Copet, Gabriel Synnaeve, Emmanuel Dupoux, and Yossi Adi. 2022 · 2022
Closest in time.
S3prl-vc: Open-source voice conversion framework with self-supervised speech representations
Wen-Chin Huang, Shu-Wen Yang, Tomoki Hayashi, Hung-Yi Lee, Shinji Watanabe, and Tomoki Toda. 2022 · 2022
Closest in time.
textless-lib: a library for textless spoken language processing
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On generative spoken language modeling from raw audio
Kushal Lakhotia, Eugene Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Abdelrahman Mohamed, et al. 2021 · 2021
Cited alongside, same era.
Voicemixer: Adversarial voice style mixup
Sang-Hoon Lee, Ji-Hoon Kim, Hyunseung Chung, and Seong-Whan Lee. 2021 · 2021
Cited alongside, same era.
Fragmentvc: Any-to-any voice conversion by end-to-end extracting and fusing fine-grained voice fragments with attention
Yist Y Lin, Chung-Ming Chien, Jheng-Hao Lin, Hung-yi Lee, and Lin-shan Lee. 2021b · 2021
Cited alongside, same era.
Any-to-many voice conversion with location-relative sequence-to-sequence modeling
Songxiang Liu, Yuewen Cao, Disong Wang, Xixin Wu, Xunying Liu, and Helen Meng. 2021 · 2021
Cited alongside, same era.
Speech resynthesis from discrete disentangled self-supervised representations
Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux. 2021 · 2021
Cited alongside, same era.
Global prosody style transfer without text transcriptions
Kaizhi Qian, Yang Zhang, Shiyu Chang, Jinjun Xiong, Chuang Gan, David Cox, and Mark Hasegawa-Johnson. 2021 · 2021
Cited alongside, same era.
SpeechBrain: A general-purpose speech toolkit
Mirco Ravanelli, Titouan Parcollet, Peter Plantinga, Aku Rouhe, Samuele Cornell, Loren Lugosch, Cem Subakan, Nauman Dawalatabad, Abdelwahab Heba, Jianyuan Zhong, Ju-Chieh Chou, Sung-Lin Yeh, Szu-Wei Fu, Chien-Feng Liao, Elena Rastorgueva, François Grondin, William Aris, Hwidong Na, Yan Gao, Renato De Mori, and Yoshua Bengio. 2021 · 2021
Cited alongside, same era.
Eugene Kharitonov, Jade Copet, Kushal Lakhotia, Tu Anh Nguyen, Paden Tomasello, Ann Lee, Ali Elkahky, Wei-Ning Hsu, Abdelrahman Mohamed, Emmanuel Dupoux, and Yossi Adi. 2022a · 2022
Closest in time.
Assem-vc: Realistic voice conversion by assembling modern speech synthesis techniques
Kang-wook Kim, Seung-won Park, Junhyeok Lee, and Myun-chul Joe. 2022 · 2022
Closest in time.
Investigation into target speaking rate adaptation for voice conversion
Michael Kuhlmann, Fritz Seebauer, Janek Ebbers, Petra Wagner, and Reinhold Haeb-Umbach. 2022 · 2022
Closest in time.
Textless speech-to-speech translation on real data
Ann Lee, Hongyu Gong, Paul-Ambroise Duquenne, Holger Schwenk, Peng-Jen Chen, Changhan Wang, Sravya Popuri, Yossi Adi, Juan Pino, Jiatao Gu, and Wei-Ning Hsu. 2022b · 2022
Closest in time.
Generative spoken dialogue language modeling
Tu Anh Nguyen, Eugene Kharitonov, Jade Copet, Yossi Adi, Wei-Ning Hsu, Ali Elkahky, Paden Tomasello, Robin Algayres, Benoit Sagot, Abdelrahman Mohamed, et al. 2022 · 2022
Closest in time.
Sravya Popuri, Peng-Jen Chen, Changhan Wang, Juan Pino, Yossi Adi, Jiatao Gu, Wei-Ning Hsu, and Ann Lee. 2022 · 2022
Closest in time.
Contentvec: An improved self-supervised speech representation by disentangling speakers
Kaizhi Qian, Yang Zhang, Heting Gao, Junrui Ni, Cheng-I Lai, David Cox, Mark Hasegawa-Johnson, and Shiyu Chang. 2022 · 2022
Closest in time.
Disentangling prosody representations with unsupervised speech reconstruction
Leyuan Qu, Taihao Li, Cornelius Weber, Theresa Pekarek-Rosin, Fuji Ren, and Stefan Wermter. 2022 · 2022
Closest in time.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2022 · 2022
Closest in time.
Emotional voice conversion: Theory, databases and esd
Kun Zhou, Berrak Sisman, Rui Liu, and Haizhou Li. 2022 · 2022
Closest in time.
Scaling speech technology to 1,000+ languages
Vineel Pratap, Andros Tjandra, Bowen Shi, Paden Tomasello, Arun Babu, Sayani Kundu, Ali Elkahky, Zhaoheng Ni, Apoorv Vyas, Maryam Fazel-Zarandi, et al. 2023 · 2023
Closest in time.