Fetching the paper…
Reading the bibliography…
People change their tones of voice, often accompanied by nonverbal vocalizations (NVs) such as laughter and cries, to convey rich emotions.
“A circumplex model of affect.,”
James A Russell, · 1980
Earlier work this paper cites.
“SWITCHBOARD: Telephone speech corpus for research and development,”
John J Godfrey, Edward C Holliman, and Jane McDaniel, · 1992
Earlier work this paper cites.
“The Fisher corpus: A resource for the next generations of speech-to-text.,”
Christopher Cieri, David Miller, and Kevin Walker, · 2004
Earlier work this paper cites.
“The AMI meeting corpus: A pre-announcement,”
Jean Carletta, Simone Ashby, Sebastien Bourban, Mike Flynn, Mael Guillemot, Thomas Hain, Jaroslav Kadlec, Vasilis Karaiskos, Wessel Kraaij, Melissa Kronenthal, et al., · 2005
Earlier work this paper cites.
“LibriSpeech: an ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Emotional end-to-end neural speech synthesizer,”
Younggun Lee, Azam Rabiee, and Soo-Young Lee, · 2017
Earlier work this paper cites.
“Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings,”
Reza Lotfian and Carlos Busso, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Neural ordinary differential equations,”
Ricky TQ Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud, · 2018
Earlier work this paper cites.
“Capturing, representing, and interacting with laughter,”
Kimiko Ryokai, Elena Durań Loṕez, Noura Howell, Jon Gillick, and David Bamman, · 2018
Earlier work this paper cites.
“The ryerson audio-visual database of emotional speech and song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in north american english,”
Steven R Livingstone and Frank A Russo, · 2018
Earlier work this paper cites.
“MelGAN: Generative adversarial networks for conditional waveform synthesis,”
Kundan Kumar, Rithesh Kumar, Thibault De Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre De Brebisson, Yoshua Bengio, and Aaron C Courville, · 2019
Earlier work this paper cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Earlier work this paper cites.
“Libri-Light: A benchmark for ASR with limited or no supervision,”
Jacob Kahn, Morgane Riviere, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, Julien Karadayi, Vitaliy Liptchinsky, Ronan Collobert, Christian Fuegen, et al., · 2020
Earlier work this paper cites.
“Controllable emotion transfer for end-to-end speech synthesis,”
Tao Li, Shan Yang, Liumeng Xue, and Lei Xie, · 2021
Cited alongside, same era.
“Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition,”
Xiong Cai, Dongyang Dai, Zhiyong Wu, Xiang Li, Jingbei Li, and Helen Meng, · 2021
Cited alongside, same era.
“Robust laughter detection in noisy environments.,”
Jon Gillick, Wesley Deng, Kimiko Ryokai, and David Bamman, · 2021
Cited alongside, same era.
“Speech synthesis with mixed emotions,”
Kun Zhou, Berrak Sisman, Rajib Rana, Björn W Schuller, and Haizhou Li, · 2022
Cited alongside, same era.
“MsEmoTTS: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis,”
Yi Lei, Shan Yang, Xinsheng Wang, and Lei Xie, · 2022
Cited alongside, same era.
“Text-driven emotional style control and cross-speaker style transfer in neural tts,”
“emotion2vec: Self-supervised pre-training for speech emotion representation,”
Ziyang Ma, Zhisheng Zheng, Jiaxin Ye, Jinchao Li, Zhifu Gao, Shiliang Zhang, and Xie Chen, · 2023
Later among the works it cites.
“Robust speech recognition via large-scale weak supervision,”
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever, · 2023
Later among the works it cites.
“Neural codec language models are zero-shot text-to-speech synthesizers,”
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al., · 2023
Later among the works it cites.
“Seamless: Multilingual expressive and streaming speech translation,”
Loïc Barrault, Yu-An Chung, Mariano Coria Meglioli, David Dale, Ning Dong, Mark Duppenthaler, Paul-Ambroise Duquenne, Brian Ellis, Hady Elsahar, Justin Haaheim, et al., · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yookyung Shin, Younggun Lee, Suhee Jo, Yeongtae Hwang, and Taesu Kim, · 2022
Cited alongside, same era.
“DNSMOS P.835: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,”
Chandan KA Reddy, Vishak Gopal, and Ross Cutler, · 2022
Cited alongside, same era.
“WavLM: Large-scale self-supervised pre-training for full stack speech processing,”
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al., · 2022
Cited alongside, same era.
“An emotion speech synthesis method based on VITS,”
Wei Zhao and Zheng Yang, · 2023
Cited alongside, same era.
“EmoDiff: Intensity controllable emotional text-to-speech with soft-label guidance,”
Yiwei Guo, Chenpeng Du, Xie Chen, and Kai Yu, · 2023
Cited alongside, same era.
“EmoMix: Emotion mixing via diffusion models for emotional speech synthesis,”
Haobin Tang, Xulong Zhang, Jianzong Wang, Ning Cheng, and Jing Xiao, · 2023
Cited alongside, same era.
“QI-TTS: Questioning intonation control for emotional speech synthesis,”
Haobin Tang, Xulong Zhang, Jianzong Wang, Ning Cheng, and Jing Xiao, · 2023
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al., · 2023
Later among the works it cites.
Haobin Tang, Xulong Zhang, Ning Cheng, Jing Xiao, and Jianzong Wang, · 2024
Closest in time.
“EmoSphere-TTS: Emotional style and intensity modeling via spherical emotion vector for controllable emotional text-to-speech,”
Deok-Hyeon Cho, Hyung-Seok Oh, Seung-Bin Kim, Sang-Hoon Lee, and Seong-Whan Lee, · 2024
Closest in time.
“VoiceBox: Text-guided multilingual universal speech generation at scale,”
Matthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer, Leda Sari, Rashel Moritz, Mary Williamson, Vimal Manohar, Yossi Adi, Jay Mahadeokar, et al., · 2024
Closest in time.
“Making flow-matching-based zero-shot text-to-speech laugh as you like,”
Naoyuki Kanda, Xiaofei Wang, Sefik Emre Eskimez, Manthan Thakker, Hemin Yang, Zirun Zhu, Min Tang, Canrun Li, Steven Tsai, Zhen Xiao, et al., · 2024
Closest in time.
“NaturalSpeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,”
Zeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan, Detai Xin, Dongchao Yang, Yanqing Liu, Yichong Leng, Kaitao Song, Siliang Tang, et al., · 2024
Closest in time.
“An investigation of noise robustness for flow-matching-based zero-shot tts,”
Xiaofei Wang, Sefik Emre Eskimez, Manthan Thakker, Hemin Yang, Zirun Zhu, Min Tang, Yufei Xia, Jinzhu Li, Sheng Zhao, Jinyu Li, et al., · 2024
Closest in time.
“JVNV: A corpus of Japanese emotional speech with verbal content and nonverbal expressions,”
Detai Xin, Junfeng Jiang, Shinnosuke Takamichi, Yuki Saito, Akiko Aizawa, and Hiroshi Saruwatari, · 2024
Closest in time.
“DiariST: Streaming speech translation with speaker diarization,”
Mu Yang, Naoyuki Kanda, Xiaofei Wang, Junkun Chen, Peidong Wang, Jian Xue, Jinyu Li, and Takuya Yoshioka, · 2024
Closest in time.
“Total-duration-aware duration modeling for text-to-speech systems,”
Sefik Emre Eskimez, Xiaofei Wang, Manthan Thakker, Chung-Hsien Tsai, Canrun Li, Zhen Xiao, Hemin Yang, Zirun Zhu, Min Tang, Jinyu Li, et al., · 2024
Closest in time.