Fetching the paper…
Reading the bibliography…
Machine recognition of an atypical speech like whispered speech, is a challenging task.
“On the calculation of filter coefficients for maximum entropy spectral analysis,”
N. Anderson, · 1978
Earlier work this paper cites.
“What’s in a whisper?,”
Tartter Vivien, · 1989
Earlier work this paper cites.
“A linear prediction algorithm in low bit rate speech coding improved by multi-band excitation model,”
Ming Yang, Fenghai Qiu, and Fuyuan Mo, · 2001
Earlier work this paper cites.
“Reconstruction of speech from whispers,”
Robert Morris and Mark Clements, · 2002
Earlier work this paper cites.
“The CHAINS corpus: Characterizing individual speakers,”
F. Cummins, Marco Grimaldi, T. Leonard, and Juraj Simko, · 2006
Earlier work this paper cites.
“Analysis-by-synthesis method for whisper-speech reconstruction,”
Farzaneh Ahmadi, Ian Vince McLoughlin, and Hamid Reza Sharifzadeh, · 2008
Earlier work this paper cites.
“Atypical Speech,”
Georg Stemmer, Elmar Nöth, and Vijay Parsa, · 2010
Earlier work this paper cites.
“The kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely, · 2011
Earlier work this paper cites.
“Statistical voice conversion techniques for body-conducted unvoiced speech enhancement,”
Tomoki Toda, Mikihiro Nakagiri, and Kiyohiro Shikano, · 2012
Earlier work this paper cites.
“Reconstruction of continuous voiced speech from whispers.,”
Ian Vince McLoughlin, Jingjie Li, and Yan Song, · 2013
Earlier work this paper cites.
“A cross-lingual adaptation approach for rapid development of speech recognizers for learning disabled users,”
Marek Bohac, Michaela Kucharova, Zoraida Callejas, Jan Nouza, and Petr Červa, · 2014
Earlier work this paper cites.
“Whisper-to-speech conversion using restricted boltzmann machine arrays,”
Jing-jie Li, Ian V McLoughlin, Li-Rong Dai, and Zhen-hua Ling, · 2014
Cited alongside, same era.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Cited alongside, same era.
“Atypical speech and voices: Corpora, classification, coaching and conversion,”
Björn Schuller, Tiago H. Falk, Vijay Parsa, and Elmar Nöth, · 2015
Cited alongside, same era.
“Reconstruction of phonated speech from whispers using formant-derived plausible pitch modulation,”
Ian V Mcloughlin, Hamid Reza Sharifzadeh, Su Lim Tan, Jingjie Li, and Yan Song, · 2015
Cited alongside, same era.
“Performance analysis of mandarin whispered speech recognition based on normal speech training model,”
Chen Xueqin, Zhao Heming, and Fan Xiaohe, · 2016
Cited alongside, same era.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Later among the works it cites.
“Whispered speech to neutral speech conversion using bidirectional lstms.,”
G Nisha Meenakshi and Prasanta Kumar Ghosh, · 2018
Later among the works it cites.
“Improved training of end-to-end attention models for speech recognition,”
Albert Zeyer, Kazuki Irie, Ralf Schlüter, and Hermann Ney, · 2018
Later among the works it cites.
“Personalizing ASR for Dysarthric and Accented Speech with Limited Data,”
Joel Shor, Dotan Emanuel, Oran Lang, Omry Tuval, Michael Brenner, Julie Cattiau, Avinatan Hassidim, and Yossi Matias, · 2019
Later among the works it cites.
“End-to-End Dysarthric Speech Recognition Using Multiple Databases,”
Yuki Takashima, Tetsuya Takiguchi, and Yasuo Ariki, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Generative modeling of pseudo-whisper for robust whispered speech recognition,”
Shabnam Ghaffarzadegan, Hynek Bořil, and John HL Hansen, · 2016
Cited alongside, same era.
“World: A vocoder-based high-quality speech synthesis system for real-time applications,”
Masanori Morise, Fumiya YOKOMORI, and Kenji Ozawa, · 2016
Cited alongside, same era.
“Phoneme-Discriminative Features for Dysarthric Speech Conversion,”
Ryo Aihara, Tetsuya Takiguchi, and Yasuo Ariki, · 2017
Cited alongside, same era.
“Whispered speech recognition using deep denoising autoencoder and inverse filtering,”
DJordje T Grozdić and Slobodan T Jovičić, · 2017
Cited alongside, same era.
“Voice conversion based on a mixture density network,”
Mohsen Ahangar, Mostafa Ghorbandoost, Sudhendu Sharma, and Mark JT Smith, · 2017
Cited alongside, same era.
“Project Euphonia: A Research initiative focused on helping people with atypical speech, https://sites.research.google/euphonia/about,”
Google Inc.,
Cited in the paper.
Hailun Lian, Yuting Hu, Weiwei Yu, Jian Zhou, and Wenming Zheng, · 2019
Later among the works it cites.
“Direct speech-to-speech translation with a sequence-to-sequence model,”
Ye Jia, Ron J. Weiss, Fadi Biadsy, Wolfgang Macherey, Melvin Johnson, Zhifeng Chen, and Yonghui Wu, · 2019
Later among the works it cites.
“Whisper augmented end-to-end/hybrid speech recognition system - cyclegan approach,”
Prithvi Raj Reddy Gudepu, Gowtham Prudhvi Vadisetti, Abhishek Niranjan, Kinnera Saranu, Raghava Sarma, M. Ali Basha Shaik, and Periyasamy Paramasivam, · 2020
Closest in time.
“librosa/librosa: 0.7.2,” Jan. 2020
Brian McFee et. al., · 2020
Closest in time.
“End-to-End Whispered Speech Recognition with Frequency-weighted Approaches and Pseudo Whisper Pre-training,”
Heng-Jui Chang, Alexander H. Liu, Hung yi Lee, and Lin-Shan Lee, · 2021
Closest in time.