Fetching the paper…
Reading the bibliography…
Generative AI has demonstrated impressive performance in various fields, among which speech synthesis is an interesting direction.
Automatic generation of control signals for a parallel formant speech synthesizer. In ICASSP’76. IEEE International Conference on Acoustics, Speech, and Signal Processing , Vol. 1. IEEE, 690–693
P Seeviour, J Holmes, and M Judd. 1976 · 1976
Earlier work this paper cites.
Rule synthesis of speech from dyadic units. In ICASSP’77. IEEE International Conference on Acoustics, Speech, and Signal Processing , Vol. 2. IEEE, 568–570
Joseph Olive. 1977 · 1977
Earlier work this paper cites.
Software for a cascade/parallel formant synthesizer
Dennis H Klatt. 1980 · 1980
Earlier work this paper cites.
Speaker adaptation through vector quantization. In ICASSP’86. IEEE International Conference on Acoustics, Speech, and Signal Processing , Vol. 11. IEEE, 2643–2646
Kiyohiro Shikano, Kai-Fu Lee, and Raj Reddy. 1986 · 1986
Earlier work this paper cites.
Review of text-to-speech conversion for English
Dennis H Klatt. 1987 · 1987
Earlier work this paper cites.
Voice conversion through vector quantization
Masanobu Abe, Satoshi Nakamura, Kiyohiro Shikano, and Hisao Kuwabara. 1990 · 1990
Earlier work this paper cites.
Pitch-synchronous waveform processing techniques for text-to-speech synthesis using diphones
Eric Moulines and Francis Charpentier. 1990 · 1990
Earlier work this paper cites.
An adaptive algorithm for mel-cepstral analysis of speech.. In icassp , Vol. 92. 137–140
Toshiaki Fukada, Keiichi Tokuda, Takao Kobayashi, and Satoshi Imai. 1992 · 1992
Earlier work this paper cites.
Mel-generalized cepstral analysis-a unified approach to speech spectral estimation.. In ICSLP , Vol. 94. 18–22
Keiichi Tokuda, Takao Kobayashi, Takashi Masuko, and Satoshi Imai. 1994 · 1994
Earlier work this paper cites.
Speech parameter generation from HMM using dynamic features. In 1995 International Conference on Acoustics, Speech, and Signal Processing , Vol. 1. IEEE, 660–663
Keiichi Tokuda, Takao Kobayashi, and Satoshi Imai. 1995 · 1995
Earlier work this paper cites.
Unit selection in a concatenative speech synthesis system using a large speech database. In 1996 IEEE International Conference on Acoustics, Speech, and Signal Processing Conference Proceedings , Vol. 1. IEEE, 373–376
Andrew J Hunt and Alan W Black. 1996 · 1996
Earlier work this paper cites.
Hidden Markov model based voice conversion using dynamic characteristics of speaker. In European Conference On Speech Communication And Technology . Eurospeech, 2519–2522
Eun-Kyoung Kim, Sangho Lee, and Yung-Hwan Oh. 1997 · 1997
Earlier work this paper cites.
Continuous probabilistic transform for voice conversion
Yannis Stylianou, Olivier Cappé, and Eric Moulines. 1998 · 1998
Earlier work this paper cites.
Restructuring speech representations using a pitch-adaptive time–frequency smoothing and an instantaneous-frequency-based F0 extraction: Possible role of a repetitive structure in sounds
Hideki Kawahara, Ikuyo Masuda-Katsuse, and Alain De Cheveigne. 1999 · 1999
Earlier work this paper cites.
Simultaneous modeling of spectrum, pitch and duration in HMM-based speech synthesis. In Sixth European Conference on Speech Communication and Technology
Takayoshi Yoshimura, Keiichi Tokuda, Takashi Masuko, Takao Kobayashi, and Tadashi Kitamura. 1999 · 1999
Earlier work this paper cites.
Speech parameter generation algorithms for HMM-based speech synthesis. In 2000 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No. 00CH37100) , Vol. 3. IEEE, 1315–1318
Keiichi Tokuda, Takayoshi Yoshimura, Takashi Masuko, Takao Kobayashi, and Tadashi Kitamura. 2000 · 2000
Earlier work this paper cites.
Marsyas: A framework for audio analysis
George Tzanetakis and Perry Cook. 2000 · 2000
Earlier work this paper cites.
Aperiodicity extraction and control using mixed mode excitation and group delay manipulation for a high quality speech analysis, modification and synthesis system STRAIGHT. In Second international workshop on models and analysis of vocal emissions for biomedical applications
Hideki Kawahara, Jo Estill, and Osamu Fujimura. 2001 · 2001
Earlier work this paper cites.
Computing mel-frequency cepstral coefficients on the power spectrum. In 2001 IEEE international conference on acoustics, speech, and signal processing. Proceedings (cat. No. 01CH37221) , Vol. 1. IEEE, 73–76
Sirko Molau, Michael Pitz, Ralf Schluter, and Hermann Ney. 2001 · 2001
Earlier work this paper cites.
Simultaneous modeling of phonetic and prosodic parameters, and characteristic conversion for HMM-based text-to-speech systems
Takayoshi Yoshimura. 2002 · 2002
Earlier work this paper cites.
Blind source separation and independent component analysis: A review
Seungjin Choi, Andrzej Cichocki, Hyung-Min Park, and Soo-Young Lee. 2005 · 2005
Earlier work this paper cites.
Temporal derivative-based spectrum and mel-cepstrum audio steganalysis
Qingzhong Liu, Andrew H Sung, and Mengyu Qiao. 2009 · 2009
Earlier work this paper cites.
Statistical parametric speech synthesis
Heiga Zen, Keiichi Tokuda, and Alan W Black. 2009 · 2009
Earlier work this paper cites.
Time-frequency analysis of acoustic signals in the audio-frequency range generated during Hadfield’s steel friction
SA Dobrynin, EA Kolubaev, A Yu Smolin, AI Dmitriev, and SG Psakhie. 2010 · 2010
Earlier work this paper cites.
Speech dereverberation based on variance-normalized delayed linear prediction
Tomohiro Nakatani, Takuya Yoshioka, Keisuke Kinoshita, Masato Miyoshi, and Biing-Hwang Juang. 2010 · 2010
Earlier work this paper cites.
Relative attributes. In 2011 International Conference on Computer Vision . IEEE, 503–510
Devi Parikh and Kristen Grauman. 2011 · 2011
Earlier work this paper cites.
Speech synthesis techniques. A survey. In International Workshop on Systems, Signal Processing and their Applications, WOSSPA . IEEE, 67–70
Youcef Tabet and Mohamed Boughazi. 2011 · 2011
Earlier work this paper cites.
Intonation conversion from neutral to expressive speech. In Twelfth Annual Conference of the International Speech Communication Association
Christophe Veaux and Xavier Rodet. 2011 · 2011
Earlier work this paper cites.
Speech synthesis based on hidden Markov models
Keiichi Tokuda, Yoshihiko Nankaku, Tomoki Toda, Heiga Zen, Junichi Yamagishi, and Keiichiro Oura. 2013 · 2013
Earlier work this paper cites.
Beyond NMF: Time-Domain Audio Source Separation without Phase Reconstruction.. In ISMIR . 369–374
Kazuyoshi Yoshii, Ryota Tomioka, Daichi Mochihashi, and Masataka Goto. 2013 · 2013
Earlier work this paper cites.
Statistical parametric speech synthesis using deep neural networks. In 2013 ieee international conference on acoustics, speech and signal processing . IEEE, 7962–7966
Heiga Ze, Andrew Senior, and Mike Schuster. 2013 · 2013
Earlier work this paper cites.
DNN-based speech bandwidth expansion and its application to adding high-frequency missing features for automatic speech recognition of narrowband speech. In Sixteenth Annual Conference of the International Speech Communication Association
Kehuang Li, Zhen Huang, Yong Xu, and Chin-Hui Lee. 2015 · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics. 2256–2265
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015 · 2015
Earlier work this paper cites.
Voice conversion using deep bidirectional long short-term memory based recurrent neural networks. In 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 4869–4873
Lifa Sun, Shiyin Kang, Kun Li, and Helen Meng. 2015 · 2015
Earlier work this paper cites.
Unidirectional long short-term memory recurrent neural network with recurrent output layer for low-latency speech synthesis. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 4470–4474
Heiga Zen and Haşim Sak. 2015 · 2015
Earlier work this paper cites.
Phone-aware LSTM-RNN for voice conversion. In 2016 IEEE 13th International Conference on Signal Processing (ICSP) . IEEE, 177–182
Jiahao Lai, Bo Chen, Tian Tan, Sibo Tong, and Kai Yu. 2016 · 2016
Earlier work this paper cites.
Voice conversion using convolutional neural networks
Shariq Mobin and Joan Bruna. 2016 · 2016
Earlier work this paper cites.
Multi-output RNN-LSTM for multiple speaker speech synthesis with a-interpolation model. In SSW9: 9th ISCA Workshop on Speech Synthesis: proceedings: Sunnyvale (CA, USA): September 13-15, 2016 . Institute of Electrical and Electronics Engineers (IEEE), 112–117
Santiago Pascual and Antonio Bonafonte Cávez. 2016 · 2016
Earlier work this paper cites.
WaveNet: A Generative Model for Raw Audio. In The 9th ISCA Speech Synthesis Workshop
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu. 2016 · 2016
Earlier work this paper cites.
Deep voice: Real-time neural text-to-speech. In International Conference on Machine Learning . PMLR, 195–204
Sercan Ö Arık, Mike Chrzanowski, Adam Coates, Gregory Diamos, Andrew Gibiansky, Yongguo Kang, Xian Li, John Miller, Andrew Ng, Jonathan Raiman, et al · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events. In 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 776–780
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter. 2017 · 2017
Earlier work this paper cites.
Deep voice 2: Multi-speaker neural text-to-speech
Andrew Gibiansky, Sercan Arik, Gregory Diamos, John Miller, Kainan Peng, Wei Ping, Jonathan Raiman, and Yanqi Zhou. 2017 · 2017
Earlier work this paper cites.
Audio super resolution using neural networks
Volodymyr Kuleshov, S Zayd Enam, and Stefano Ermon. 2017 · 2017
Earlier work this paper cites.
SEGAN: Speech enhancement generative adversarial network
Santiago Pascual, Antonio Bonafonte, and Joan Serra. 2017 · 2017
Earlier work this paper cites.
Deep Voice 3: 2000-Speaker Neural Text-to-Speech
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan Ömer Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller. 2017 · 2017
Earlier work this paper cites.
Speech Enhancement Using Bayesian Wavenet.. In Interspeech . 2013–2017
Kaizhi Qian, Yang Zhang, Shiyu Chang, Xuesong Yang, Dinei Florêncio, and Mark Hasegawa-Johnson. 2017 · 2017
Earlier work this paper cites.
Char2wav: End-to-end speech synthesis
Jose Sotelo, Soroush Mehri, Kundan Kumar, Joao Felipe Santos, Kyle Kastner, Aaron Courville, and Yoshua Bengio. 2017 · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Cited alongside, same era.
Tacotron: Towards end-to-end speech synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al · 2017
Cited alongside, same era.
End-to-end waveform utterance enhancement for direct evaluation metrics optimization by fully convolutional neural networks
Szu-Wei Fu, Tao-Wei Wang, Yu Tsao, Xugang Lu, and Hisashi Kawai. 2018 · 2018
Cited alongside, same era.
DNN-based source enhancement to increase objective sound quality assessment score
Yuma Koizumi, Kenta Niwa, Yusuke Hioka, Kazunori Kobayashi, and Yoichi Haneda. 2018 · 2018
Cited alongside, same era.
Time-frequency networks for audio super-resolution. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 646–650
Teck Yian Lim, Raymond A Yeh, Yijia Xu, Minh N Do, and Mark Hasegawa-Johnson. 2018 · 2018
Upsampling artifacts in neural audio synthesis. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 3005–3009
Jordi Pons, Santiago Pascual, Giulio Cengarle, and Joan Serrà. 2021 · 2021
Later among the works it cites.
Grad-tts: A diffusion probabilistic model for text-to-speech. 8599–8608
Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, and Mikhail Kudinov. 2021 · 2021
Later among the works it cites.
Simon Rouard and Gaëtan Hadjeres. 2021 · 2021
Later among the works it cites.
A flow-based neural network for time domain speech enhancement. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 5754–5758
Martin Strauss and Bernd Edler. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Model-based STFT phase recovery for audio source separation
Paul Magron, Roland Badeau, and Bertrand David. 2018 · 2018
Cited alongside, same era.
Parallel wavenet: Fast high-fidelity speech synthesis. In International conference on machine learning . PMLR, 3918–3926
Aaron Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George Driessche, Edward Lockhart, Luis Cobo, Florian Stimberg, et al · 2018
Cited alongside, same era.
Clarinet: Parallel wave generation in end-to-end text-to-speech
Wei Ping, Kainan Peng, and Jitong Chen. 2018 · 2018
Cited alongside, same era.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions. In 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 4779–4783
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al · 2018
Cited alongside, same era.
Time-frequency masking-based speech enhancement using generative adversarial network. In 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 5039–5043
Meet H Soni, Neil Shah, and Hemant A Patil. 2018 · 2018
Cited alongside, same era.
High fidelity speech synthesis with adversarial networks
Mikołaj Bińkowski, Jeff Donahue, Sander Dieleman, Aidan Clark, Erich Elsen, Norman Casagrande, Luis C Cobo, and Karen Simonyan. 2019 · 2019
Cited alongside, same era.
Temporal FiLM: Capturing Long-Range Sequence Dependencies with Feature-Wise Modulations
Sawyer Birnbaum, Volodymyr Kuleshov, Zayd Enam, Pang Wei W Koh, and Stefano Ermon. 2019 · 2019
Cited alongside, same era.
Xu Tan, Tao Qin, Frank Soong, and Tie-Yan Liu. 2021 · 2021
Later among the works it cites.
Wave-tacotron: Spectrogram-free end-to-end text-to-speech synthesis. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 5679–5683
Ron J Weiss, RJ Skerry-Ryan, Eric Battenberg, Soroosh Mariooryad, and Diederik P Kingma. 2021 · 2021
Later among the works it cites.
Tackling the generative learning trilemma with denoising diffusion gans
Zhisheng Xiao, Karsten Kreis, and Arash Vahdat. 2021 · 2021
Later among the works it cites.
Restoring degraded speech via a modified diffusion model
Jianwei Zhang, Suren Jayasuriya, and Visar Berisha. 2021 · 2021
Later among the works it cites.
Cold diffusion: Inverting arbitrary image transforms without noise
Arpit Bansal, Eitan Borgnia, Hong-Min Chu, Jie S Li, Hamid Kazemi, Furong Huang, Micah Goldblum, Jonas Geiping, and Tom Goldstein. 2022 · 2022
Later among the works it cites.
A survey on generative diffusion model
Hanqun Cao, Cheng Tan, Zhangyang Gao, Guangyong Chen, Pheng-Ann Heng, and Stan Z Li. 2022 · 2022
Later among the works it cites.
Infergrad: Improving Diffusion Models for Vocoder by Considering Inference in Training. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 8432–8436
Zehua Chen, Xu Tan, Ke Wang, Shifeng Pan, Danilo Mandic, Lei He, and Sheng Zhao. 2022 · 2022
Later among the works it cites.
EmoDiff: Intensity Controllable Emotional Text-to-Speech with Soft-Label Guidance
Yiwei Guo, Chenpeng Du, Xie Chen, and Kai Yu. 2022 · 2022
Later among the works it cites.
NU-Wave 2: A general neural audio upsampling model for various sampling rates
Seungu Han and Junhyeok Lee. 2022 · 2022
Later among the works it cites.
FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis
Rongjie Huang, Max WY Lam, Jun Wang, Dan Su, Dong Yu, Yi Ren, and Zhou Zhao. 2022a · 2022
Later among the works it cites.
GenerSpeech: Towards Style Transfer for Generalizable Out-Of-Domain Text-to-Speech Synthesis
Rongjie Huang, Yi Ren, Jinglin Liu, Chenye Cui, and Zhou Zhao. 2022b · 2022
Later among the works it cites.
ProDiff: Progressive Fast Diffusion Model For High-Quality Text-to-Speech
Rongjie Huang, Zhou Zhao, Huadai Liu, Jinglin Liu, Chenye Cui, and Yi Ren. 2022c · 2022
Later among the works it cites.
Any-speaker Adaptive Text-To-Speech Synthesis with Diffusion Models
Minki Kang, Dongchan Min, and Sung Ju Hwang. 2022 · 2022
Later among the works it cites.
Denoising diffusion restoration models
Bahjat Kawar, Michael Elad, Stefano Ermon, and Jiaming Song. 2022 · 2022
Later among the works it cites.
Guided-TTS 2: A Diffusion Model for High-quality Adaptive Text-to-Speech with Untranscribed Data
Sungwon Kim, Heeseung Kim, and Sungroh Yoon. 2022 · 2022
Later among the works it cites.
WaveFit: An Iterative and Non-autoregressive Neural Vocoder based on Fixed-Point Iteration
Yuma Koizumi, Kohei Yatabe, Heiga Zen, and Michiel Bacchiani. 2022a · 2022
Later among the works it cites.
SpecGrad: Diffusion Probabilistic Model based Neural Vocoder with Adaptive Noise Spectral Shaping
Yuma Koizumi, Heiga Zen, Kohei Yatabe, Nanxin Chen, and Michiel Bacchiani. 2022b · 2022
Later among the works it cites.
BDDM: Bilateral Denoising Diffusion Models for Fast and High-Quality Speech Synthesis
Max WY Lam, Jun Wang, Dan Su, and Dong Yu. 2022 · 2022
Later among the works it cites.
Zero-Shot Voice Conditioning for Denoising Diffusion TTS Models
Alon Levkovitch, Eliya Nachmani, and Lior Wolf. 2022 · 2022
Later among the works it cites.
DiffGAN-TTS: High-Fidelity and Efficient Text-to-Speech with Denoising Diffusion GANs
Songxiang Liu, Dan Su, and Dong Yu. 2022 · 2022
Later among the works it cites.
Conditional diffusion probabilistic model for speech enhancement. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 7402–7406
Yen-Ju Lu, Zhong-Qiu Wang, Shinji Watanabe, Alexander Richard, Cheng Yu, and Yu Tsao. 2022 · 2022
Later among the works it cites.
Solving Audio Inverse Problems with a Diffusion Model
Eloi Moliner, Jaakko Lehtinen, and Vesa Välimäki. 2022 · 2022
Later among the works it cites.
TUNet: A Block-online Bandwidth Extension Model based on Transformers and Self-supervised Pretraining. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 161–165
Viet-Anh Nguyen, Anh HT Nguyen, and Andy WH Khong. 2022 · 2022
Later among the works it cites.
Full-band General Audio Synthesis with Score-based Diffusion
Santiago Pascual, Gautam Bhattacharya, Chunghsin Yeh, Jordi Pons, and Joan Serrà. 2022 · 2022
Later among the works it cites.
Speech enhancement and dereverberation with diffusion-based generative models
Julius Richter, Simon Welker, Jean-Marie Lemercier, Bunlong Lay, and Timo Gerkmann. 2022 · 2022
Later among the works it cites.
Unsupervised vocal dereverberation with diffusion-based generative models
Koichi Saito, Naoki Murata, Toshimitsu Uesaka, Chieh-Hsin Lai, Yuhta Takida, Takao Fukui, and Yuki Mitsufuji. 2022 · 2022
Later among the works it cites.
A Versatile Diffusion-based Generative Refiner for Speech Enhancement
Ryosuke Sawata, Naoki Murata, Yuhta Takida, Toshimitsu Uesaka, Takashi Shibuya, Shusuke Takahashi, and Yuki Mitsufuji. 2022 · 2022
Later among the works it cites.
Diffusion-based Generative Speech Source Separation
Robin Scheibler, Youna Ji, Soo-Whan Chung, Jaeuk Byun, Soyeon Choe, and Min-Seok Choi. 2022 · 2022
Later among the works it cites.
Universal Speech Enhancement with Score-based Diffusion
Joan Serrà, Santiago Pascual, Jordi Pons, R Oguz Araz, and Davide Scaini. 2022 · 2022
Later among the works it cites.
ITÔN: End-to-end audio generation with Itô stochastic differential equations
Ziqiang Shi and Shoule Wu. 2022 · 2022
Later among the works it cites.
Speech Enhancement with Score-Based Generative Models in the Complex STFT Domain
Simon Welker, Julius Richter, and Timo Gerkmann. 2022 · 2022
Later among the works it cites.
ItôWave: Itô Stochastic Differential Equation is all You Need for Wave Generation. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 8422–8426
Shoule Wu and Ziqiang Shi. 2022 · 2022
Later among the works it cites.
NoreSpeech: Knowledge Distillation based Conditional Diffusion Model for Noise-robust Expressive TTS
Dongchao Yang, Songxiang Liu, Jianwei Yu, Helin Wang, Chao Weng, and Yuexian Zou. 2022a · 2022
Later among the works it cites.
Diffsound: Discrete Diffusion Model for Text-to-sound Generation
Dongchao Yang, Jianwei Yu, Helin Wang, Wen Wang, Chao Weng, Yuexian Zou, and Dong Yu. 2022c · 2022
Later among the works it cites.
Diffusion probabilistic modeling for video generation
Ruihan Yang, Prakhar Srivastava, and Stephan Mandt. 2022b · 2022
Later among the works it cites.
Cold Diffusion for Speech Enhancement
Hao Yen, François G Germain, Gordon Wichern, and Jonathan Le Roux. 2022 · 2022
Later among the works it cites.
Conditioning and Sampling in Variational Diffusion Models for Speech Super-resolution
Chin-Yun Yu, Sung-Lin Yeh, György Fazekas, and Hao Tang. 2022 · 2022
Later among the works it cites.
One Small Step for Generative AI, One Giant Leap for AGI: A Complete Survey on ChatGPT in AIGC Era
Chaoning Zhang, Chenshuang Zhang, Chenghao Li, Sheng Zheng, Yu Qiao, Sumit Kumar Dam, Mengchun Zhang, Jung Uk Kim, Seong Tae Kim, Gyeong-Moon Park, Jinwoo Choi, Sung-Ho Bae, Lik-Hang Lee, Pan Hui, In So Kweon, and Choong Seon Hong. 2023b · 2023
Closest in time.
Text-to-image Diffusion Models in Generative AI: A Survey
Chenshuang Zhang, Chaoning Zhang, Mengchun Zhang, and In So Kweon. 2023c · 2023
Closest in time.
A Complete Survey on Generative AI (AIGC): Is ChatGPT from GPT-4 to GPT-5 All You Need?
Chaoning Zhang, Chenshuang Zhang, Sheng Zheng, Yu Qiao, Chenghao Li, Mengchun Zhang, Sumit Kumar Dam, Chu Myaet Thwal, Ye Lin Tun, Le Luang Huy, Donguk kim, Sung-Ho Bae, Lik-Hang Lee, Yang Yang, Heng Tao Shen, In So Kweon, and Choong Seon Hong. 2023d · 2023
Closest in time.
A Survey on Graph Diffusion Models: Generative AI in Science for Molecule, Protein and Material
Mengchun Zhang, Maryam Qamar, Taegoo Kang, Yuna Jung, Chenshuang Zhang, Sung-Ho Bae, and Chaoning Zhang. 2023a · 2023
Closest in time.
Metricgan: Generative adversarial networks based black-box metric scores optimization for speech enhancement. In International Conference on Machine Learning . PMLR, 2031–2041
Szu-Wei Fu, Chien-Feng Liao, Yu Tsao, and Shou-De Lin. 2019 · 2041
Closest in time.