Fetching the paper…
Reading the bibliography…
Modifying the pitch and timing of an audio signal are fundamental audio editing operations with applications in speech manipulation, audio-visual synchronization, and singing voice editing and synthesis.
“Tentative standards for sound level meters,”
RG McCurdy, · 1936
Earlier work this paper cites.
“The viterbi algorithm,”
G David Forney, · 1973
Earlier work this paper cites.
“Linear prediction: A tutorial review,”
John Makhoul, · 1975
Earlier work this paper cites.
“On the relation between pitch excursion size and prominence,”
Antonius CM Rietveld and C Gussenhovent, · 1985
Earlier work this paper cites.
“Dither in digital audio,”
John Vanderkooy and Stanley P Lipshitz, · 1987
Earlier work this paper cites.
“Pitch-synchronous waveform processing techniques for text-to-speech synthesis using diphones,”
E. Moulines and F. Charpentier, · 1990
Earlier work this paper cites.
“Discrete-time speech signal processing: Principles and practice,”
T. Quatieri, · 2001
Earlier work this paper cites.
“YIN, a fundamental frequency estimator for speech and music,”
Alain De Cheveigné and Hideki Kawahara, · 2002
Earlier work this paper cites.
“Implementation of realtime straight speech manipulation system: Report on its first implementation,”
H. Banno, H. Hata, M. Morise, T. Takahashi, T. Irino, and H. Kawahara, · 2007
Earlier work this paper cites.
“Can we automatically transform speech recorded on common consumer devices in real-world environments into professional production quality speech?—a dataset, insights, and challenges,”
Gautham J Mysore, · 2014
Earlier work this paper cites.
“A simple limiter in python,” https://gist.github.com/bastibe/747283c55aad66404046 , 2015
B. Bechtold, · 2015
Earlier work this paper cites.
“Algorithms to measure audio programme loudness and true-peak audio level,”
International Telecommunications Union, · 2015
Earlier work this paper cites.
“WORLD: a vocoder-based high-quality speech synthesis system for real-time applications,”
M. Morise, F. Yokomori, and K. Ozawa, · 2016
Cited alongside, same era.
“Sound quality comparison among high-quality vocoders by using re-synthesized speech,”
Masanori Morise and Yusuke Watanabe, · 2018
Cited alongside, same era.
“A hybrid dsp/deep learning approach to real-time full-band speech enhancement,”
Jean-Marc Valin, · 2018
Cited alongside, same era.
“CREPE: A convolutional representation for pitch estimation,”
Jong Wook Kim, Justin Salamon, Peter Li, and Juan Pablo Bello, · 2018
Cited alongside, same era.
“The ryerson audio-visual database of emotional speech and song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in north american english,”
Steven R Livingstone and Frank A Russo, · 2018
Cited alongside, same era.
“On the convergence of adam and beyond,”
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar, · 2019
Later among the works it cites.
“Controllable neural prosody synthesis,”
M. Morrison, Z. Jin, J. Salamon, N. J. Bryan, and G. J. Mysore, · 2020
Later among the works it cites.
“Hider-finder-combiner: An adversarial architecture for general speech signal modification,”
Jacob J Webber, Olivier Perrotin, and Simon King, · 2020
Later among the works it cites.
“Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,”
Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim, · 2020
Later among the works it cites.
“LPCNet: DSP-boosted neural speech synthesis,” https://jmvalin.ca/demo/lpcnet/ ,
J.-M. Valin, · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Waveglow: A flow-based generative network for speech synthesis,”
R. Prenger, R. Valle, and B. Catanzaro, · 2019
Cited alongside, same era.
“Neural source-filter-based waveform model for statistical parametric speech synthesis,”
Xin Wang, Shinji Takaki, and Junichi Yamagishi, · 2019
Cited alongside, same era.
“LPCNet: Improving neural speech synthesis through linear prediction,”
J.-M. Valin and J. Skoglund, · 2019
Cited alongside, same era.
“High quality, lightweight and adaptable tts using lpcnet,”
Zvi Kons, Slava Shechtman, Alex Sorin, Carmel Rabinovitz, and Ron Hoory, · 2019
Cited alongside, same era.
“LPCNet,” https://github.com/mozilla/LPCNet , 2019
J.-M. Valin, · 2019
Cited alongside, same era.
“CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92),”
Junichi Yamagishi, Christophe Veaux, Kirsten MacDonald, et al., · 2019
Cited alongside, same era.
“torchcrepe,” https://github.com/maxrmorrison/torchcrepe , 2020
M. Morrison, · 2020
Later among the works it cites.
“psola,” https://github.com/maxrmorrison/psola , 2020
M. Morrison, · 2020
Later among the works it cites.
“Hifi-gan: High-fidelity denoising and dereverberation based on speech deep features in adversarial networks,”
Jiaqi Su, Zeyu Jin, and Adam Finkelstein, · 2020
Later among the works it cites.
“Quasi-periodic parallel wavegan: A non-autoregressive raw waveform generative model with pitch-dependent dilated convolution neural network,”
Yi-Chiao Wu, Tomoki Hayashi, Takuma Okamoto, Hisashi Kawai, and Tomoki Toda, · 2021
Closest in time.
Reo Yoneyama, Yi-Chiao Wu, and Tomoki Toda, · 2021
Closest in time.
“pyworld,” https://github.com/JeremyCCHsu/Python-Wrapper-for-World-Vocoder , 2021
J. Hsu, · 2021
Closest in time.