Fetching the paper…
Reading the bibliography…
We present VoiceShop, a novel speech-to-speech framework that can modify multiple attributes of speech, such as age, gender, accent, and speech style, in a single forward pass while preserving the input speaker's timbre.
Acoustic Theory of Speech Production: With Calculations Based on X-Ray Studies of Russian Articulations
G. Fant · 1971
Earlier work this paper cites.
Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
Solving Ordinary Differential Equations I: Nonstiff Problems
E. Hairer, S.P. Nørsett, and G. Wanner · 2008
Earlier work this paper cites.
Visualizing Data using t-SNE
Laurens van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
DiffWave: A Versatile Diffusion Model for Audio Synthesis
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro · 2009
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Sequence to Sequence Learning with Neural Networks, 2014
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le · 2014
Earlier work this paper cites.
A Neural Algorithm of Artistic Style, 2015
Leon A. Gatys, Alexander S. Ecker, and Matthias Bethge · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
U-Net: Convolutional Networks for Biomedical Image Segmentation, 2015
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Domain-Adversarial Training of Neural Networks, 2016
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor Lempitsky · 2016
Earlier work this paper cites.
Variational Inference with Normalizing Flows, 2016
Danilo Jimenez Rezende and Shakir Mohamed · 2016
Earlier work this paper cites.
WaveNet: A Generative Model for Raw Audio, 2016
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit
Christophe Veaux, Junichi Yamagishi, Kirsten MacDonald, et al · 2016
Earlier work this paper cites.
One Model To Learn Them All, 2017
Lukasz Kaiser, Aidan N. Gomez, Noam Shazeer, Ashish Vaswani, Niki Parmar, Llion Jones, and Jakob Uszkoreit · 2017
Earlier work this paper cites.
Attention is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Neural Ordinary Differential Equations
Ricky TQ Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud · 2018
Earlier work this paper cites.
Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al · 2018
Earlier work this paper cites.
Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis, 2018
Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ Skerry-Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Fei Ren, Ye Jia, and Rif A. Saurous · 2018
Earlier work this paper cites.
ESPnet: End-to-End Speech Processing Toolkit
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, Adithya Renduchintala, and Tsubasa Ochiai · 2018
Earlier work this paper cites.
Accent Conversion Using Phonetic Posteriorgrams
Guanlong Zhao, Sinem Sonsaat, John Levis, Evgeny Chukharev-Hudilainen, and Ricardo Gutierrez-Osuna · 2018
Earlier work this paper cites.
Common Voice: A Massively-Multilingual Speech Corpus
Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M Tyers, and Gregor Weber · 2019
Earlier work this paper cites.
Decoupled Weight Decay Regularization, 2019
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le · 2019
Earlier work this paper cites.
AUTOVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss, 2019
Kaizhi Qian, Yang Zhang, Shiyu Chang, Xuesong Yang, and Mark Hasegawa-Johnson · 2019
Earlier work this paper cites.
Foreign Accent Conversion by Synthesizing Speech from Phonetic Posteriorgrams
Guanlong Zhao, Shaojin Ding, and Ricardo Gutierrez-Osuna · 2019
Cited alongside, same era.
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations, 2020
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Cited alongside, same era.
Location-Relative Attention Mechanisms For Robust Long-Form Speech Synthesis
Eric Battenberg, RJ Skerry-Ryan, Soroosh Mariooryad, Daisy Stanton, David Kao, Matt Shannon, and Tom Bagby · 2020
Cited alongside, same era.
Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck · 2020
Cited alongside, same era.
Conformer: Convolution-Augmented Transformer for Speech Recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang · 2020
Cited alongside, same era.
MuLan: A Joint Embedding of Music Audio and Natural Language
Qingqing Huang, Aren Jansen, Joonseok Lee, Ravi Ganti, Judith Yue Li, and Daniel PW Ellis · 2022
Later among the works it cites.
CopyCat2: A Single Model for Multi-Speaker TTS and Many-to-Many Fine-Grained Prosody Transfer, 2022
Sri Karlapati, Penny Karanasou, Mateusz Lajszczak, Ammar Abbas, Alexis Moinet, Peter Makarov, Ray Li, Arent van Korlaar, Simon Slangen, and Thomas Drugman · 2022
Later among the works it cites.
Towards Disentangled Speech Representations, 2022
Cal Peyser, Ronny Huang Andrew Rosenberg Tara N. Sainath, Michael Picheny, and Kyunghyun Cho · 2022
Later among the works it cites.
Progressive Distillation for Fast Sampling of Diffusion Models, 2022
Tim Salimans and Jonathan Ho · 2022
Later among the works it cites.
Denoising Diffusion Implicit Models, 2022
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Denoising Diffusion Probabilistic Models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
TFGAN: Time and Frequency Domain Based Generative Adversarial Network for High-Fidelity Speech Synthesis, 2020
Qiao Tian, Yi Chen, Zewang Zhang, Heng Lu, Linghui Chen, Lei Xie, and Shan Liu · 2020
Cited alongside, same era.
StyleFlow: Attribute-Conditioned Exploration of StyleGAN-Generated Images using Conditional Continuous Normalizing Flows
Rameen Abdal, Peihao Zhu, Niloy J Mitra, and Peter Wonka · 2021
Cited alongside, same era.
SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network, 2021
William Chan, Daniel Park, Chris Lee, Yu Zhang, Quoc Le, and Mohammad Norouzi · 2021
Cited alongside, same era.
Accelerating Continuous Normalizing Flow with Trajectory Polynomial Regularization
Han-Hsien Huang and Mi-Yen Yeh · 2021
Cited alongside, same era.
Perceiver: General Perception with Iterative Attention
Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Joao Carreira · 2021
Cited alongside, same era.
Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech, 2021
Jaehyeon Kim, Jungil Kong, and Juhee Son · 2021
Cited alongside, same era.
Yongmao Zhang, Zhichao Wang, Peiji Yang, Hongshen Sun, Zhisheng Wang, and Lei Xie · 2022
Later among the works it cites.
Multilingual Multiaccented Multispeaker TTS with RADTTS, 2023
Rohan Badlani, Rafael Valle, Kevin J. Shih, João Felipe Santos, Siddharth Gururani, and Bryan Catanzaro · 2023
Later among the works it cites.
Invisible Watermarking for Audio Generation Diffusion Models, 2023
Xirong Cao, Xiang Li, Divyesh Jadav, Yanzhao Wu, Zhehui Chen, Chen Zeng, and Wenqi Wei · 2023
Later among the works it cites.
DDDM-VC: Decoupled Denoising Diffusion Models with Disentangled Representation and Prior Mixup for Verified Robust Voice Conversion, 2023
Ha-Yeong Choi, Sang-Hoon Lee, and Seong-Whan Lee · 2023
Later among the works it cites.
Interpretable Style Transfer for Text-to-Speech with ControlVAE and Diffusion Bridge, 2023
Wenhao Guan, Tao Li, Yishuang Li, Hukai Huang, Qingyang Hong, and Lin Li · 2023
Later among the works it cites.
Zero-Shot Accent Conversion using Pseudo Siamese Disentanglement Network, 2023
Dongya Jia, Qiao Tian, Kainan Peng, Jiaxin Li, Yuanzhe Chen, Mingbo Ma, Yuping Wang, and Yuxuan Wang · 2023
Later among the works it cites.
Voice-Preserving Zero-Shot Multiple Accent Conversion, 2023
Mumin Jin, Prashant Serai, Jilong Wu, Andros Tjandra, Vimal Manohar, and Qing He · 2023
Later among the works it cites.
UnitSpeech: Speaker-Adaptive Speech Synthesis with Untranscribed Data, 2023
Heeseung Kim, Sungwon Kim, Jiheum Yeom, and Sungroh Yoon · 2023
Later among the works it cites.
Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale
Matthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer, Leda Sari, Rashel Moritz, Mary Williamson, Vimal Manohar, Yossi Adi, Jay Mahadeokar, et al · 2023
Later among the works it cites.
Multimodal Foundation Models: From Specialists to General-Purpose Assistants
Chunyuan Li, Zhe Gan, Zhengyuan Yang, Jianwei Yang, Linjie Li, Lijuan Wang, and Jianfeng Gao · 2023
Later among the works it cites.
DINO-VITS: Data-Efficient Noise-Robust Zero-Shot Voice Cloning via Multi-Tasking with Self-Supervised Speaker Verification Loss, 2023
Vikentii Pankov, Valeria Pronina, Alexander Kuzmin, Maksim Borisov, Nikita Usoltsev, Xingshan Zeng, Alexander Golubkov, Nikolai Ermolenko, Aleksandra Shirshova, and Yulia Matveeva · 2023
Later among the works it cites.
Moûsai: Text-to-Music Generation with Long-Context Latent Diffusion, 2023
Flavio Schneider, Ojasv Kamal, Zhijing Jin, and Bernhard Schölkopf · 2023
Later among the works it cites.
Modelling Low-Resource Accents without Accent-Specific TTS Frontend, 2023
Georgi Tinchev, Marta Czarnowska, Kamil Deja, Kayoko Yanagisawa, and Marius Cotescu · 2023
Later among the works it cites.
Accented Text-to-Speech Synthesis with Limited Data, 2023
Xuehao Zhou, Mingyang Zhang, Yi Zhou, Zhizheng Wu, and Haizhou Li · 2023
Later among the works it cites.
WavMark: Watermarking for Audio Generation, 2024
Guangyu Chen, Yu Wu, Shujie Liu, Tao Liu, Xiaoyong Du, and Furu Wei · 2024
Closest in time.
Audio Deepfake Detection with Self-Supervised WavLM and Multi-Fusion Attentive Classifier, 2024
Yinlin Guo, Haofan Huang, Xi Chen, He Zhao, and Yuehai Wang · 2024
Closest in time.
Collaborative Watermarking for Adversarial Speech Synthesis, 2024
Lauri Juvela and Xin Wang · 2024
Closest in time.
Proactive Detection of Voice Cloning with Localized Watermarking, 2024
Robin San Roman, Pierre Fernandez, Alexandre Défossez, Teddy Furon, Tuan Tran, and Hady Elsahar · 2024
Closest in time.
Diffusion Models: A Comprehensive Survey of Methods and Applications, 2024
Ling Yang, Zhilong Zhang, Yang Song, Shenda Hong, Runsheng Xu, Yue Zhao, Wentao Zhang, Bin Cui, and Ming-Hsuan Yang · 2024
Closest in time.
SingFake: Singing Voice Deepfake Detection, 2024
Yongyi Zang, You Zhang, Mojtaba Heydari, and Zhiyao Duan · 2024
Closest in time.