Fetching the paper…
Reading the bibliography…
Voice-enabled interactions provide more human-like experiences in many popular IoT systems.
Voice conversion: Factors responsible for quality. In ICASSP’85. IEEE International Conference on Acoustics, Speech, and Signal Processing , Vol. 10. IEEE, 748–751
D Childers, B Yegnanarayana, and Ke Wu. 1985 · 1985
Earlier work this paper cites.
Inferring speakers’ physical attributes from their voices
Robert M Krauss, Robin Freyberg, and Ezequiel Morsella. 2002 · 2002
Earlier work this paper cites.
A system for transforming the emotion in speech: Combining data-driven conversion techniques for prosody and voice quality. In Eighth Annual Conference of the International Speech Communication Association
Zeynep Inanoglu and Steve Young. 2007 · 2007
Earlier work this paper cites.
Using linguistic cues for the automatic recognition of personality in conversation and text
François Mairesse, Marilyn A Walker, Matthias R Mehl, and Roger K Moore. 2007 · 2007
Earlier work this paper cites.
Fast and reliable F0 estimation method based on the period extraction of vocal fold vibration of singing voice and speech. In Audio Engineering Society Conference: 35th International Conference: Audio for Games
Masanori Morise, Hideki Kawahara, and Haruhiro Katayose. 2009 · 2009
Earlier work this paper cites.
Estimation of unknown speaker’s height from speech
Iosif Mporas and Todor Ganchev. 2009 · 2009
Earlier work this paper cites.
Adaptations in humans for assessing physical strength from the voice
Aaron Sell, Gregory A Bryant, Leda Cosmides, John Tooby, Daniel Sznycer, Christopher Von Rueden, Andre Krauss, and Michael Gurven. 2010 · 2010
Earlier work this paper cites.
Computational paralinguistics: emotion, affect and personality in speech and language processing
Björn Schuller and Anton Batliner. 2013 · 2013
Earlier work this paper cites.
The INTERSPEECH 2013 computational paralinguistics challenge: Social signals, conflict, emotion, autism
Björn Schuller, Stefan Steidl, Anton Batliner, Alessandro Vinciarelli, Klaus Scherer, Fabien Ringeval, Mohamed Chetouani, Felix Weninger, Florian Eyben, Erik Marchi, et al · 2013
Earlier work this paper cites.
Generative adversarial nets. In Advances in neural information processing systems . 2672–2680
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Regulating the internet of things: first steps toward managing discrimination, privacy, security and consent
Scott R Peppet. 2014 · 2014
Cited alongside, same era.
CheapTrick, a spectral envelope estimator for high-quality speech synthesis
Masanori Morise. 2015 · 2015
Cited alongside, same era.
Spoofing and countermeasures for speaker verification: A survey
Zhizheng Wu, Nicholas Evans, Tomi Kinnunen, Junichi Yamagishi, Federico Alegre, and Haizhou Li. 2015 · 2015
Cited alongside, same era.
D4C, a band-aperiodicity estimator for high-quality speech synthesis
Masanori Morise. 2016 · 2016
Cited alongside, same era.
WORLD: a vocoder-based high-quality speech synthesis system for real-time applications
Masanori Morise, Fumiya Yokomori, and Kenji Ozawa. 2016 · 2016
Cited alongside, same era.
VoxCeleb2: Deep speaker recognition
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman. 2018a · 2018
Later among the works it cites.
Voxceleb2: Deep speaker recognition
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman. 2018b · 2018
Later among the works it cites.
The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in North American English
Steven R Livingstone and Frank A Russo. 2018 · 2018
Later among the works it cites.
Hidebehind: Enjoy Voice Input with Voiceprint Unclonability and Anonymity. In Proceedings of the 16th ACM Conference on Embedded Networked Sensor Systems . ACM, 82–94
Jianwei Qian, Haohua Du, Jiahui Hou, Linlin Chen, Taeho Jung, and Xiang-Yang Li. 2018 · 2018
Later among the works it cites.
Emotional expression in psychiatric conditions: New technology for clinicians
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adieu features? end-to-end speech emotion recognition using a deep convolutional recurrent network. In 2016 IEEE international conference on acoustics, speech and signal processing (ICASSP) . 5200–5204
George Trigeorgis, Fabien Ringeval, Raymond Brueckner, Erik Marchi, Mihalis A Nicolaou, Björn Schuller, and Stefanos Zafeiriou. 2016 · 2016
Cited alongside, same era.
Monkey says, monkey does: security and privacy on voice assistants
Efthimios Alepis and Constantinos Patsakis. 2017 · 2017
Cited alongside, same era.
Speaking Style Conversion from Normal to Lombard Speech Using a Glottal Vocoder and Bayesian GMMs. 1363–1367
Ana Ramírez López, Shreyas Seshadri, Lauri Juvela, Okko Räsänen, and Paavo Alku. 2017 · 2017
Cited alongside, same era.
Unpaired image-to-image translation using cycle-consistent adversarial networks. In Proceedings of the IEEE international conference on computer vision . 2223–2232
Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros. 2017 · 2017
Cited alongside, same era.
Cloud Speech-to-Text - Speech Recognition | Cloud Speech-to-Text | Google Cloud
[n. d.]
Cited in the paper.
Karol Grabowski, Agnieszka Rynkiewicz, Amandine Lassalle, Simon Baron-Cohen, Björn Schuller, Nicholas Cummins, Alice Baird, Justyna Podgórska-Bednarz, Agata Pieniążek, and Izabela Łucka. 2019 · 2019
Closest in time.
CycleGAN-VC2: Improved CycleGAN-based Non-parallel Voice Conversion. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, and Nobukatsu Hojo. 2019 · 2019
Closest in time.
marcogdepinto/Emotion-Classification-Ravdess
Marcogdepinto. 2019 · 2019
Closest in time.
Utterance-level Aggregation For Speaker Recognition In The Wild
Weidi Xie, Arsha Nagrani, Joon Son Chung, and Andrew Zisserman. 2019 · 2019
Closest in time.
Multi-task self-supervised visual learning. In Proceedings of the IEEE International Conference on Computer Vision . 2051–2060
Carl Doersch and Andrew Zisserman. 2017 · 2060
Closest in time.