Fetching the paper…
Reading the bibliography…
Speech emotion conversion is the task of modifying the perceived emotion of a speech utterance while preserving the lexical content and speaker identity.
Decoding of inconsistent communications
Albert Mehrabian and Morton Wiener. 1967 · 1967
Earlier work this paper cites.
An argument for basic emotions
Paul Ekman. 1992 · 1992
Earlier work this paper cites.
Yet another algorithm for pitch tracking
K. Kasi and S. A. Zahorian. 2002 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni et al. 2002 · 2002
Earlier work this paper cites.
Transforming spectrum and prosody for emotional voice conversion with non-parallel training data
Kun Zhou, Berrak Sisman, and Haizhou Li. 2020a · 2002
Earlier work this paper cites.
The cmu arctic speech databases
John Kominek and Alan W Black. 2004 · 2004
Earlier work this paper cites.
Multi-target emotional voice conversion with neural vocoders
Songxiang Liu, Yuewen Cao, and Helen Meng. 2020 · 2004
Earlier work this paper cites.
Fastspeech 2: Fast and high-quality end-to-end text to speech
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2020 · 2006
Earlier work this paper cites.
Data augmenting contrastive learning of speech representations in the time domain
Eugene Kharitonov, Morgane Rivière, Gabriel Synnaeve, Lior Wolf, Pierre-Emmanuel Mazaré, Matthijs Douze, and Emmanuel Dupoux. 2021b · 2007
Earlier work this paper cites.
Ravi Shankar, Hsi-Wei Hsieh, Nicolas Charon, and Archana Venkataraman. 2020 · 2007
Earlier work this paper cites.
Hierarchical multi-grained generative model for expressive speech synthesis
Yukiya Hono, Kazuna Tsuboi, Kei Sawada, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, and Keiichi Tokuda. 2020 · 2009
Earlier work this paper cites.
Data-driven emotion conversion in spoken english
Zeynep Inanoglu and Steve Young. 2009 · 2009
Earlier work this paper cites.
CROWDMOS: An approach for crowdsourcing mean opinion score studies
F. Ribeiro, D. Florêncio, C. Zhang, and M. Seltzer. 2011 · 2011
Earlier work this paper cites.
Gmm-based emotional voice conversion using spectrum and prosody features
Ryo Aihara, Ryoichi Takashima, Tetsuya Takiguchi, and Yasuo Ariki. 2012 · 2012
Earlier work this paper cites.
The ilsp/innoetics text-to-speech system for the blizzard challenge 2013
Aimilios Chalamandaris, Pirros Tsiakoulis, Sotiris Karabetsos, Spyros Raptis, and I LTD. 2013 · 2013
Earlier work this paper cites.
Exemplar-based emotional voice conversion using non-negative matrix factorization
Ryo Aihara, Reina Ueda, Tetsuya Takiguchi, and Yasuo Ariki. 2014 · 2014
Earlier work this paper cites.
LibriSpeech: an ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
End-to-end text-dependent speaker verification
Georg Heigold, Ignacio Moreno, Samy Bengio, and Noam Shazeer. 2016 · 2016
Earlier work this paper cites.
Emotional voice conversion using deep neural networks with mcc and f0 features
Zhaojie Luo, Tetsuya Takiguchi, and Yasuo Ariki. 2016 · 2016
Earlier work this paper cites.
Deep bidirectional lstm modeling of timbre and prosody for emotional voice conversion
Huaiping Ming, Dong-Yan Huang, Lei Xie, Jie Wu, Minghui Dong, and Haizhou Li. 2016 · 2016
Earlier work this paper cites.
Variational inference for acoustic unit discovery
Lucas Ondel, Lukáš Burget, and Jan Černockỳ. 2016 · 2016
Earlier work this paper cites.
Hidden Markov Model variational autoencoder for acoustic unit discovery
Janek Ebbers, Jahn Heymann, Lukas Drude, Thomas Glarner, Reinhold Haeb-Umbach, and Bhiksha Raj. 2017 · 2017
Earlier work this paper cites.
Unsupervised learning of disentangled and interpretable representations from sequential data
Wei-Ning Hsu, Yu Zhang, and James Glass. 2017 · 2017
Cited alongside, same era.
Neural discrete representation learning
Aaron van den Oord, Oriol Vinyals, et al. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
The emotional voices database: Towards controlling the emotion dimension in voice generation systems
Adaeze Adigwe, Noé Tits, Kevin El Haddad, Sarah Ostadabbas, and Thierry Dutoit. 2018 · 2018
Cited alongside, same era.
Nonparallel emotional speech conversion
Jian Gao, Deep Chakraborty, Hamidou Tembine, and Olaitan Olaleye. 2018 · 2018
HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020 · 2020
Later among the works it cites.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
Unsupervised speech decomposition via triple information bottleneck
Kaizhi Qian, Yang Zhang, Shiyu Chang, Mark Hasegawa-Johnson, and David Cox. 2020 · 2020
Later among the works it cites.
Multi-task self-supervised learning for robust speech recognition
M. Ravanelli, J. Zhong, S. Pascual, P. Swietojanski, J. Monteiro, J. Trmal, and Y. Bengio. 2020 · 2020
Later among the works it cites.
Stargan for emotional speech conversion: Validated by data augmentation of end-to-end emotion recognition
Georgios Rizos, Alice Baird, Max Elliott, and Björn Schuller. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Full Bayesian Hidden Markov Model variational autoencoder for acoustic unit discovery
Thomas Glarner, Patrick Hanebrink, Janek Ebbers, and Reinhold Haeb-Umbach. 2018 · 2018
Cited alongside, same era.
Crepe: A convolutional representation for pitch estimation
Jong Wook Kim, Justin Salamon, Peter Li, and Juan Pablo Bello. 2018 · 2018
Cited alongside, same era.
Investigating different representations for modeling and controlling multiple emotions in dnn-based speech synthesis
Jaime Lorenzo-Trueba, Gustav Eje Henter, Shinji Takaki, Junichi Yamagishi, Yosuke Morino, and Yuta Ochiai. 2018 · 2018
Cited alongside, same era.
Natural TTS synthesis by conditioning WaveNet on MEL spectrogram predictions
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, R. A. Saurous, Y. Agiomvrgiannakis, and Y. Wu. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin et al. 2019 · 2019
Cited alongside, same era.
Fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Cited alongside, same era.
Sequence-to-sequence modelling of f0 for speech emotion conversion
Carl Robinson, Nicolas Obin, and Axel Roebel. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Principal style components: Expressive style control and cross-speaker transfer in neural tts
Alexander Sorin, Slava Shechtman, and Ron Hoory. 2020 · 2020
Later among the works it cites.
Unsupervised pre-training of bidirectional speech encoders via masked reconstruction
W. Wang, Q. Tang, and K. Livescu. 2020 · 2020
Later among the works it cites.
Self-supervised representations improve end-to-end speech translation
Anne Wu, Changhan Wang, Juan Pino, and Jiatao Gu. 2020 · 2020
Later among the works it cites.
Audio ALBERT: A lite BERT for self-supervised learning of audio representation
Po-Han Chi, Pei-Hung Chung, Tsung-Han Wu, Chun-Cheng Hsieh, Shang-Wen Li, and Hung-yi Lee. 2021 · 2021
Closest in time.
Sequence-to-sequence emotional voice conversion with strength control
Heejin Choi and Minsoo Hahn. 2021 · 2021
Closest in time.
HuBERT: How much can a bad teacher benefit ASR pre-training?
Wei-Ning Hsu, Yao-Hung Hubert Tsai, Benjamin Bolte, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021 · 2021
Closest in time.
Expressive text-to-speech using style tag
Minchan Kim, Sung Jun Cheon, Byoung Jin Choi, Jong Jin Kim, and Nam Soo Kim. 2021 · 2021
Closest in time.
Semi-supervised spoken language understanding via self-supervised speech and language model pretraining
Cheng-I Lai, Yung-Sung Chuang, Hung-Yi Lee, Shang-Wen Li, and James Glass. 2021 · 2021
Closest in time.
Generative spoken language modeling from raw audio
Kushal Lakhotia, Evgeny Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Adelrahman Mohamed, et al. 2021 · 2021
Closest in time.
Do people agree on how positive emotions are expressed? a survey of four emotions and five modalities across 11 cultures
Kunalan Manokara, Mirna Đurić, Agneta Fischer, and Disa Sauter. 2021 · 2021
Closest in time.
Speech resynthesis from discrete disentangled self-supervised representations
Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux. 2021 · 2021
Closest in time.
Global rhythm style transfer without text transcriptions
Kaizhi Qian, Yang Zhang, Shiyu Chang, Jinjun Xiong, Chuang Gan, David Cox, and Mark Hasegawa-Johnson. 2021 · 2021
Closest in time.
A survey on neural speech synthesis
Xu Tan, Tao Qin, Frank Soong, and Tie-Yan Liu. 2021 · 2021
Closest in time.
Superb: Speech processing universal performance benchmark
Shu-wen Yang et al. 2021 · 2021
Closest in time.
Vaw-gan for disentanglement and recomposition of emotional elements in speech
Kun Zhou, Berrak Sisman, and Haizhou Li. 2021b · 2021
Closest in time.
Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset
Kun Zhou, Berrak Sisman, Rui Liu, and Haizhou Li. 2021d · 2021
Closest in time.