Fetching the paper…
Reading the bibliography…
The zero-shot text-to-speech (TTS) method, based on speaker embeddings extracted from reference speech using self-supervised learning (SSL) speech representations, can reproduce speaker characteristics very accurately.
“Catastrophic forgetting in connectionist networks,”
Robert M French, · 1999
Earlier work this paper cites.
“Intonation pattern and duration differences in imitated speech,”
Elisabeth Zetterholm, · 2002
Earlier work this paper cites.
“MUSAN: A music, speech, and noise corpus,” arXiv, 2015,
David Snyder, Guoguo Chen, and Daniel Povey, · 2015
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2015
Earlier work this paper cites.
“End-to-end text-dependent speaker verification,”
Georg Heigold, Ignacio Moreno, Samy Bengio, and Noam Shazeer, · 2016
Earlier work this paper cites.
“Speaker adaptation in DNN-based speech synthesis using d-vectors.,”
Rama Doddipatla, Norbert Braunschweiler, and Ranniery Maia, · 2017
Earlier work this paper cites.
“Deep speaker embeddings for short-duration speaker verification.,”
Gautam Bhattacharya, Md Jahangir Alam, and Patrick Kenny, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, et al., · 2017
Earlier work this paper cites.
“Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, et al., · 2018
Earlier work this paper cites.
“X-vectors: Robust DNN embeddings for speaker recognition,”
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur, · 2018
Earlier work this paper cites.
“Soft-target training with ambiguous emotional utterances for DNN-based speech emotion classification,”
Atsushi Ando, Satoshi Kobashikawa, Hosana Kamiyama, Ryo Masumura, Yusuke Ijima, and Yushi Aono, · 2018
Earlier work this paper cites.
“Parameter-efficient transfer learning for NLP,”
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly, · 2019
Earlier work this paper cites.
“Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation,”
Yi Luo and Nima Mesgarani, · 2019
Cited alongside, same era.
“FastSpeech 2: Fast and high-quality end-to-end text to speech,”
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2020
Cited alongside, same era.
“Zero-shot multi-speaker text-to-speech with state-of-the-art neural speaker embeddings,”
Erica Cooper, Cheng-I Lai, Yusuke Yasuda, Fuming Fang, Xin Wang, Nanxin Chen, and Junichi Yamagishi, · 2020
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Cited alongside, same era.
“HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,”
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, · 2020
Cited alongside, same era.
“WavLM: Large-scale self-supervised pre-training for full stack speech processing,”
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al., · 2022
Later among the works it cites.
“Boosting self-supervised embeddings for speech enhancement,”
Kuo-Hsuan Hung, Szu wei Fu, Huan-Hsin Tseng, Hsin-Tien Chiang, Yu Tsao, and Chii-Wann Lin, · 2022
Later among the works it cites.
“End-to-End integration of speech recognition, speech enhancement, and self-supervised learning representation,”
Xuankai Chang, Takashi Maekaku, Yuya Fujita, and Shinji Watanabe, · 2022
Later among the works it cites.
“Why does self-supervised learning for speech recognition benefit speaker recognition?,”
Sanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu, Zhuo Chen, Peidong Wang, et al., · 2022
Later among the works it cites.
“Exploring efficient-tuning methods in self-supervised speech models,”
Zih-Ching Chen, Chin-Lun Fu, Chih-Ying Liu, Shang-Wen Daniel Li, and Hung-yi Lee, · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Investigating on incorporating pretrained and learnable speaker representations for multi-speaker multi-style text-to-speech,”
Chung-Ming Chien, Jheng-Hao Lin, Chien-yu Huang, Po-chun Hsu, and Hung-yi Lee, · 2021
Cited alongside, same era.
“Phoneme duration modeling using speech rhythm-based speaker embeddings for multi-speaker speech synthesis,”
Kenichi Fujita, Atsushi Ando, and Yusuke Ijima, · 2021
Cited alongside, same era.
“HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Cited alongside, same era.
“DenoiSpeech: Denoising text to speech with frame-level noise modeling,”
Chen Zhang, Yi Ren, Xu Tan, Jinglin Liu, Kejun Zhang, Tao Qin, Sheng Zhao, and Tie-Yan Liu, · 2021
Cited alongside, same era.
“Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,”
Jaehyeon Kim, Jungil Kong, and Juhee Son, · 2021
Cited alongside, same era.
“DRSpeech: Degradation-robust text-to-speech synthesis with frame-level and utterance-level acoustic representation learning,”
Takaaki Saeki, Kentaro Tachibana, and Ryuichi Yamamoto, · 2022
Cited alongside, same era.
“Fine-grained noise control for multispeaker speech synthesis,”
Karolos Nikitaras, Georgios Vamvoukakis, Nikolaos Ellinas, Konstantinos Klapsas, Konstantinos Markopoulos, Spyros Raptis, June Sig Sung, Gunu Jho, Aimilios Chalamandaris, and Pirros Tsiakoulis, · 2022
Cited alongside, same era.
“Flamingo: a visual language model for few-shot learning,”
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al., · 2022
Later among the works it cites.
“JETS: Jointly Training FastSpeech2 and HiFi-GAN for End to End Text to Speech,”
Dan Lim, Sunghee Jung, and Eesung Kim, · 2022
Later among the works it cites.
“Neural codec language models are zero-shot text to speech synthesizers,” arXiv, 2023,
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, Lei He, Sheng Zhao, and Furu Wei, · 2023
Later among the works it cites.
“Zero-shot text-to-speech synthesis conditioned using self-supervised speech representation model,”
Kenichi Fujita, Takanori Ashihara, Hiroki Kanagawa, Takafumi Moriya, and Yusuke Ijima, · 2023
Later among the works it cites.
“NoreSpeech: Knowledge distillation based conditional diffusion model for noise-robust expressive TTS,”
Dongchao Yang, Songxiang Liu, Helin Wang, Jianwei Yu, Chao Weng, and Yuexian Zou, · 2023
Later among the works it cites.
“CHAPTER: Exploiting convolutional neural network adapters for self-supervised speech models,”
Zih-Ching Chen, Yu-Shun Sung, and Hung-Yi Lee, · 2023
Later among the works it cites.
“Downstream task agnostic speech enhancement with self-supervised representation loss,”
Hiroshi Sato, Ryo Masumura, Tsubasa Ochiai, Marc Delcroix, Takafumi Moriya, Takanori Ashihara, Kentaro Shinayama, Saki Mizuno, Mana Ihori, Tomohiro Tanaka, and Nobukatsu Hojo, · 2023
Later among the works it cites.