Fetching the paper…
Reading the bibliography…
Audio editing is applicable for various purposes, such as adding background sound effects, replacing a musical instrument, and repairing damaged audio.
Voice conversion through vector quantization
Masanobu Abe, Satoshi Nakamura, Kiyohiro Shikano, and Hisao Kuwabara · 1990
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Esc: Dataset for environmental sound classification
Karol J Piczak · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
An overview of voice conversion systems
Seyed Hamidreza Mohammadi and Alexander Kain · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Efficient neural audio synthesis
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron Oord, Sander Dieleman, and Koray Kavukcuoglu · 2018
Earlier work this paper cites.
Singing voice conversion with non-parallel data
Xin Chen, Wei Chu, Jinxi Guo, and Ning Xu · 2019
Earlier work this paper cites.
Audiocaps: Generating captions for audios in the wild
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim · 2019
Earlier work this paper cites.
Play as you like: Timbre-enhanced multi-modal music style transfer
Chien-Yu Lu, Min-Xin Xue, Chia-Che Chang, Che-Rung Lee, and Li Su · 2019
Earlier work this paper cites.
Melgan-vc: Voice conversion and audio style transfer on arbitrarily long samples using spectrograms
Marco Pasini · 2019
Earlier work this paper cites.
Vision-infused deep audio inpainting
Hang Zhou, Ziwei Liu, Xudong Xu, Ping Luo, and Xiaogang Wang · 2019
Earlier work this paper cites.
Groove2groove: one-shot music style transfer with supervision from synthetic data
Ondřej Cífka, Umut Şimşekli, and Gaël Richard · 2020
Cited alongside, same era.
Audio inpainting with generative adversarial network
Pirmin P Ebner and Amr Eltelt · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
Text-to-text pre-training for data-to-text tasks
Mihir Kale and Abhinav Rastogi · 2020
Cited alongside, same era.
Gacela: A generative adversarial context encoder for long audio inpainting of music
Andrés Marafioti, Piotr Majdak, Nicki Holighaus, and Nathanaël Perraudin · 2020
Cited alongside, same era.
Cmgan: Conformer-based metric-gan for monaural speech enhancement
Sherif Abdulatif, Ruizhe Cao, and Bin Yang · 2022
Later among the works it cites.
Blended diffusion for text-driven editing of natural images
Omri Avrahami, Dani Lischinski, and Ohad Fried · 2022
Later among the works it cites.
Speechpainter: Text-conditioned speech inpainting
Zalán Borsos, Matt Sharifi, and Marco Tagliasacchi · 2022
Later among the works it cites.
Instructpix2pix: Learning to follow image editing instructions
Tim Brooks, Aleksander Holynski, and Alexei A Efros · 2022
Later among the works it cites.
Vector quantized diffusion model for text-to-image synthesis
Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ondřej Mokrỳ and Pavel Rajmic · 2020
Cited alongside, same era.
Waveflow: A compact flow-based model for raw audio
Wei Ping, Kainan Peng, Kexin Zhao, and Zhao Song · 2020
Cited alongside, same era.
Unsupervised cross-domain singing voice conversion
Adam Polyak, Lior Wolf, Yossi Adi, and Yaniv Taigman · 2020
Cited alongside, same era.
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole · 2020
Cited alongside, same era.
Ilvr: Conditioning method for denoising diffusion probabilistic models
Jooyoung Choi, Sungwon Kim, Yonghyun Jeong, Youngjune Gwon, and Sungroh Yoon · 2021
Cited alongside, same era.
Fsd50k: an open dataset of human-labeled sound events
Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, and Xavier Serra · 2021
Cited alongside, same era.
Diffsvc: A diffusion probabilistic model for singing voice conversion
Songxiang Liu, Yuewen Cao, Dan Su, and Helen Meng · 2021
Cited alongside, same era.
Prompt-to-prompt image editing with cross attention control
Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or · 2022
Later among the works it cites.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans · 2022
Later among the works it cites.
Audiogen: Textually guided audio generation
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi · 2022
Later among the works it cites.
Repaint: Inpainting using denoising diffusion probabilistic models
Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Later among the works it cites.
Palette: Image-to-image diffusion models
Chitwan Saharia, William Chan, Huiwen Chang, Chris Lee, Jonathan Ho, Tim Salimans, David Fleet, and Mohammad Norouzi · 2022
Later among the works it cites.
Diffsound: Discrete diffusion model for text-to-sound generation
Dongchao Yang, Jianwei Yu, Helin Wang, Wen Wang, Chao Weng, Yuexian Zou, and Dong Yu · 2022
Later among the works it cites.
Musiclm: Generating music from text
Andrea Agostinelli, Timo I Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, et al · 2023
Closest in time.
Audioldm: Text-to-audio generation with latent diffusion models
Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, and Mark D Plumbley · 2023
Closest in time.
Chong Mou, Xintao Wang, Liangbin Xie, Jian Zhang, Zhongang Qi, Ying Shan, and Xiaohu Qie · 2023
Closest in time.
Adding conditional control to text-to-image diffusion models
Lvmin Zhang and Maneesh Agrawala · 2023
Closest in time.