Fetching the paper…
Reading the bibliography…
Novel text-to-speech systems can generate entirely new voices that were not seen during training.
“IEEE recommended practice for speech quality measurements,” IEEE, Tech. Rep., 1969, iSBN: 9781504402743. [Online]. Available: https://ieeexplore.ieee.org/document/7405210/
1969
Earlier work this paper cites.
J. L. Fitch and A. Holbrook, “Modal vocal fundamental frequency of young adults,” Archives of Otolaryngology - Head and Neck Surgery , vol. 92, no. 4, pp. 379–382, Oct. 1970. [Online]. Available: https://doi.org/10.1001/archotol.1970.04310040067012
1970
Earlier work this paper cites.
P. Ekman, “An argument for basic emotions,” Cognition and Emotion , vol. 6, no. 3-4, pp. 169–200, May 1992. [Online]. Available: https://doi.org/10.1080/02699939208411068
1992
Earlier work this paper cites.
2010
Earlier work this paper cites.
2010
Earlier work this paper cites.
C. Sofer, R. Dotsch, D. H. J. Wigboldus, and A. Todorov, “What is typical is good,” Psychological Science , vol. 26, no. 1, pp. 39–47, Dec. 2014. [Online]. Available: https://doi.org/10.1177/0956797614554955
2014
Earlier work this paper cites.
R. E. Jack and P. G. Schyns, “The human face as a dynamic tool for social communication,” Current Biology , vol. 25, no. 14, pp. R621–R634, Jul. 2015. [Online]. Available: https://doi.org/10.1016/j.cub.2015.05.052
2015
Earlier work this paper cites.
M. Walker and T. Vetter, “Changing the personality of a face: Perceived big two and big five personality factors modeled in real photographs.” Journal of Personality and Social Psychology , vol. 110, no. 4, pp. 609–624, 2016. [Online]. Available: https://doi.org/10.1037/pspp0000064
2016
Earlier work this paper cites.
K. J. Woods, M. H. Siegel, J. Traer, and J. H. McDermott, “Headphone screening to facilitate web-based auditory experiments,” Attention, Perception, & Psychophysics , vol. 79, no. 7, pp. 2064–2072, 2017
2017
Earlier work this paper cites.
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, P. Nguyen, R. Pang, I. Lopez Moreno, Y. Wu et al. , “Transfer learning from speaker verification to multispeaker text-to-speech synthesis,” Advances in neural information processing systems , vol. 31, 2018
2018
Cited alongside, same era.
R. Angulu, J. R. Tapamo, and A. O. Adewumi, “Age estimation via face images: a survey,” EURASIP Journal on Image and Video Processing , vol. 2018, no. 1, Jun. 2018. [Online]. Available: https://doi.org/10.1186/s13640-018-0278-6
2018
Cited alongside, same era.
Y. Wang, D. Stanton, Y. Zhang, R.-S. Ryan, E. Battenberg, J. Shor, Y. Xiao, Y. Jia, F. Ren, and R. A. Saurous, “Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,” in International Conference on Machine Learning . PMLR, 2018, pp. 5180–5189
2018
Cited alongside, same era.
S. R. Livingstone and F. A. Russo, “The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in North American English,” PLOS ONE , vol. 13, no. 5, p. e0196391, 2018
K. R. Prajwal, R. Mukhopadhyay, V. P. Namboodiri, and C. Jawahar, “A lip sync expert is all you need for speech to lip generation in the wild,” in Proceedings of the 28th ACM International Conference on Multimedia . ACM, Oct. 2020. [Online]. Available: https://doi.org/10.1145/3394171.3413532
2020
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
J. Kim, J. Kong, and J. Son, “Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,” in International Conference on Machine Learning . PMLR, 2021, pp. 5530–5540
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in CVPR , 2018
2018
Cited alongside, same era.
H. Ritschel, I. Aslan, S. Mertes, A. Seiderer, and E. André, “Personalized synthesis of intentional and emotional non-verbal sounds for social robots,” in 2019 8th International conference on affective computing and intelligent interaction (ACII) . IEEE, 2019, pp. 1–7
2019
Cited alongside, same era.
P. Harrison, R. Marjieh, F. Adolfi, P. van Rijn, M. Anglada-Tort, O. Tchernichovski, P. Larrouy-Maestri, and N. Jacoby, “Gibbs sampling with people,” Advances in Neural Information Processing Systems , vol. 33, pp. 10 659–10 671, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. M. Tyers, and G. Weber, “Common voice: A massively-multilingual speech corpus,” in LREC , 2020, pp. 4218–4222
2020
Cited alongside, same era.
2021
Later among the works it cites.
M. Bernard, “Phonemizer,” https://github.com/bootphon/phonemizer, 2021
2021
Later among the works it cites.
J. Kim, J. Kong, and J. Son, “Vits implementation,” https://github.com/jaywalnut310/vits, 2022
2022
Closest in time.
Sato, “Ai gahaku,” https://ai-art.tokyo, 2022
2022
Closest in time.
“Dallinger,” https://github.com/Dallinger/Dallinger, 2022
2022
Closest in time.