Fetching the paper…
Reading the bibliography…
Recent advancements in neural vocoding are predominantly driven by Generative Adversarial Networks (GANs) operating in the time-domain.
Remaking speech
Homer Dudley · 1939
Earlier work this paper cites.
The unimportance of phase in speech enhancement
Dequan Wang and Jae Lim · 1982
Earlier work this paper cites.
Signal estimation from modified short-time fourier transform
Daniel Griffin and Jae Lim · 1984
Earlier work this paper cites.
Pitch-synchronous waveform processing techniques for text-to-speech synthesis using diphones
Eric Moulines and Francis Charpentier · 1990
Earlier work this paper cites.
Unit selection in a concatenative speech synthesis system using a large speech database
Andrew J Hunt and Alan W Black · 1996
Earlier work this paper cites.
Restructuring speech representations using a pitch-adaptive time–frequency smoothing and an instantaneous-frequency-based f0 extraction: Possible role of a repetitive structure in sounds
Hideki Kawahara, Ikuyo Masuda-Katsuse, and Alain De Cheveigne · 1999
Earlier work this paper cites.
Simultaneous modeling of spectrum, pitch and duration in hmm-based speech synthesis
Takayoshi Yoshimura, Keiichi Tokuda, Takashi Masuko, Takao Kobayashi, and Tadashi Kitamura · 1999
Earlier work this paper cites.
Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs
Antony W Rix, John G Beerends, Michael P Hollier, and Andries P Hekstra · 2001
Earlier work this paper cites.
Introduction to digital audio coding and standards , volume 721
Marina Bosi and Richard E Goldberg · 2002
Earlier work this paper cites.
Perceptual phase quantization of speech
Doh-Suk Kim · 2003
Earlier work this paper cites.
Modified discrete cosine transform: Its implications for audio coding and error concealment
Ye Wang and Mikka Vilermo · 2003
Earlier work this paper cites.
The importance of phase in speech enhancement
Kuldip Paliwal, Kamil Wójcicki, and Benjamin Shannon · 2011
Earlier work this paper cites.
Perceptual importance of the phase related information in speech
Ibon Saratxaga, Inma Hernaez, Michael Pucher, Eva Navas, and Iñaki Sainz · 2012
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Daps (device and produced speech) dataset, May 2014
Gautham J. Mysore · 2014
Earlier work this paper cites.
Complex ratio masking for monaural speech separation
Donald S Williamson, Yuxuan Wang, and DeLiang Wang · 2015
Earlier work this paper cites.
Samplernn: An unconditional end-to-end neural audio generation model
Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron Courville, and Yoshua Bengio · 2016
Earlier work this paper cites.
World: a vocoder-based high-quality speech synthesis system for real-time applications
Masanori Morise, Fumiya Yokomori, and Kenji Ozawa · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Arbitrary style transfer in real-time with adaptive instance normalization
Xun Huang and Serge Belongie · 2017
Earlier work this paper cites.
Musdb18-a corpus for music separation
Zafar Rafii, Antoine Liutkus, Fabian-Robert Stöter, Stylianos Ioannis Mimilakis, and Rachel Bittner · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Chris Donahue, Julian McAuley, and Miller Puckette · 2018
Cited alongside, same era.
Gansynth: Adversarial neural audio synthesis
Jesse Engel, Kumar Krishna Agrawal, Shuo Chen, Ishaan Gulrajani, Chris Donahue, and Adam Roberts · 2018
Cited alongside, same era.
Efficient neural audio synthesis
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron Oord, Sander Dieleman, and Koray Kavukcuoglu · 2018
Cited alongside, same era.
Rave: A variational autoencoder for fast and high-quality neural audio synthesis
Antoine Caillon and Philippe Esling · 2021
Later among the works it cites.
Won Jang, Dan Lim, Jaesam Yoon, Bongwan Kim, and Juntae Kim · 2021
Later among the works it cites.
Alias-free generative adversarial networks
Tero Karras, Miika Aittala, Samuli Laine, Erik Härkönen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila · 2021
Later among the works it cites.
Chunked autoregressive gan for conditional waveform synthesis
Max Morrison, Rithesh Kumar, Kundan Kumar, Prem Seetharaman, Aaron Courville, and Yoshua Bengio · 2021
Later among the works it cites.
Speech resynthesis from discrete disentangled self-supervised representations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Takeru Miyato and Masanori Koyama · 2018
Cited alongside, same era.
Parallel wavenet: Fast high-fidelity speech synthesis
Aaron Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George Driessche, Edward Lockhart, Luis Cobo, Florian Stimberg, et al · 2018
Cited alongside, same era.
Generative adversarial network-based approach to signal reconstruction from magnitude spectrogram
Keisuke Oyamada, Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka, Nobukatsu Hojo, and Hiroyasu Ando · 2018
Cited alongside, same era.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al · 2018
Cited alongside, same era.
Phasenet: Discretized phase modeling with deep neural networks for audio source separation
Naoya Takahashi, Purvi Agrawal, Nabarun Goswami, and Yuki Mitsufuji · 2018
Cited alongside, same era.
Phase reconstruction from amplitude spectrograms based on von-mises-distribution deep neural network
Shinnosuke Takamichi, Yuki Saito, Norihiro Takamune, Daichi Kitamura, and Hiroshi Saruwatari · 2018
Cited alongside, same era.
Melgan: Generative adversarial networks for conditional waveform synthesis
Kundan Kumar, Rithesh Kumar, Thibault De Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brébisson, Yoshua Bengio, and Aaron C Courville · 2019
Cited alongside, same era.
Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux · 2021
Later among the works it cites.
Gan vocoder: Multi-resolution discriminator is all you need
Jaeseong You, Dalhyun Kim, Gyuhyeon Nam, Geumbyeol Hwang, and Gyeongsu Chae · 2021
Later among the works it cites.
Soundstream: An end-to-end neural audio codec
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi · 2021
Later among the works it cites.
Avocodo: Generative adversarial network for artifact-free vocoder
Taejun Bak, Junmo Lee, Hanbin Bae, Jinhyeok Yang, Jae-Sung Bae, and Young-Sun Joo · 2022
Later among the works it cites.
Audiolm: a language modeling approach to audio generation
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Olivier Teboul, David Grangier, Marco Tagliasacchi, and Neil Zeghidour · 2022
Later among the works it cites.
High fidelity neural audio compression
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi · 2022
Later among the works it cites.
Icassp 2022 deep noise suppression challenge
Harishchandra Dubey, Vishak Gopal, Ross Cutler, Ashkan Aazami, Sergiy Matusevych, Sebastian Braun, Sefik Emre Eskimez, Manthan Thakker, Takuya Yoshioka, Hannes Gamper, et al · 2022
Later among the works it cites.
istftnet: Fast and lightweight mel-spectrogram vocoder incorporating inverse short-time fourier transform
Takuhiro Kaneko, Kou Tanaka, Hirokazu Kameoka, and Shogo Seki · 2022
Later among the works it cites.
Bigvgan: A universal neural vocoder with large-scale training
Sang-gil Lee, Wei Ping, Boris Ginsburg, Bryan Catanzaro, and Sungroh Yoon · 2022
Later among the works it cites.
A convnet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie · 2022
Later among the works it cites.
Musika! fast infinite waveform music generation
Marco Pasini and Jan Schlüter · 2022
Later among the works it cites.
Utmos: Utokyo-sarulab system for voicemos challenge 2022
Takaaki Saeki, Detai Xin, Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, and Hiroshi Saruwatari · 2022
Later among the works it cites.
WavThruVec: Latent speech representation as intermediate features for neural speech synthesis
Hubert Siuzdak, Piotr Dura, Pol van Rijn, and Nori Jacoby · 2022
Later among the works it cites.
Mastering diverse domains through world models
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap · 2023
Closest in time.
Bark: Text-prompted generative audio model
Suno AI · 2023
Closest in time.