Fetching the paper…
Reading the bibliography…
Recent breakthroughs in zero-shot voice synthesis have enabled imitating a speaker's voice using just a few seconds of recording while maintaining a high level of realism.
“On a class of error correcting binary group codes,”
Raj Chandra Bose and Dwijendra K Ray-Chaudhuri, · 1960
Earlier work this paper cites.
“Echo hiding,”
Daniel Gruhl, Anthony Lu, and Walter Bender, · 1996
Earlier work this paper cites.
“Techniques for data hiding,”
Walter Bender, Daniel Gruhl, Norishige Morimoto, and Anthony Lu, · 1996
Earlier work this paper cites.
“Digital watermarks for audio signals,”
Laurence Boney, Ahmed H Tewfik, and Khaled N Hamdy, · 1996
Earlier work this paper cites.
“Secure spread spectrum watermarking for multimedia,”
Ingemar J Cox, Joe Kilian, F Thomson Leighton, and Talal Shamoon, · 1997
Earlier work this paper cites.
“Secure spread spectrum watermarking for multimedia,”
Ingemar J Cox, Joe Kilian, F Thomson Leighton, and Talal Shamoon, · 1997
Earlier work this paper cites.
“Audio watermarking: features, applications and algorithms,”
M. Arnold, · 2000
Earlier work this paper cites.
“Quantization index modulation: A class of provably good methods for digital watermarking and information embedding,”
Brian Chen and Gregory W Wornell, · 2001
Earlier work this paper cites.
“Wavelet-based audio watermarking techniques: robustness and fast synchronization,”
Hong Oh Kim, Bae Keun Lee, and Nam-Yong Lee, · 2001
Earlier work this paper cites.
“Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,”
Antony W Rix, John G Beerends, Michael P Hollier, and Andries P Hekstra, · 2001
Earlier work this paper cites.
“A blind audio watermarking algorithm with self-synchronization,”
Jiwu Huang, Yong Wang, and Yun Q Shi, · 2002
Earlier work this paper cites.
“Modified patchwork algorithm: A novel audio watermarking scheme,”
In-Kwon Yeo and Hyoung Joong Kim, · 2003
Earlier work this paper cites.
“Increasing robustness of lsb audio steganography using a novel embedding method,”
Nedeljko Cvejic and Tapio Seppanen, · 2004
Earlier work this paper cites.
“A novel synchronization invariant audio watermarking scheme based on dwt and dct,”
Xiang-Yang Wang and Hong Zhao, · 2006
Earlier work this paper cites.
“Nice: Non-linear independent components estimation,”
Laurent Dinh, David Krueger, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
Digital Audio Watermarking
Martin Steinebach, · 2015
Cited alongside, same era.
“Density estimation using real nvp,”
Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio, · 2016
Cited alongside, same era.
“Fma: A dataset for music analysis,”
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, and Xavier Bresson, · 2016
Cited alongside, same era.
“Audio set: An ontology and human-labeled dataset for audio events,”
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter, · 2017
Cited alongside, same era.
“Glow: Generative flow with invertible 1x1 convolutions,”
Durk P Kingma and Prafulla Dhariwal, · 2018
Cited alongside, same era.
“Robust invertible image steganography,”
Youmin Xu, Chong Mou, Yujie Hu, Jingfen Xie, and Jian Zhang, · 2022
Later among the works it cites.
“Towards blind watermarking: Combining invertible and non-invertible mechanisms,”
Rui Ma, Mengxi Guo, Yi Hou, Fan Yang, Yuan Li, Huizhu Jia, and Xiaodong Xie, · 2022
Later among the works it cites.
“Enhancing image rescaling using dual latent variables in invertible neural network,”
Min Zhang, Zhihong Pan, Xin Zhou, and C-C Jay Kuo, · 2022
Later among the works it cites.
“Neural codec language models are zero-shot text to speech synthesizers,”
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al., · 2023
Closest in time.
“Speak foreign languages with your own voice: Cross-lingual neural codec language modeling,”
Ziqiang Zhang, Long Zhou, Chengyi Wang, Sanyuan Chen, Yu Wu, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, Lei He, Sheng Zhao, and Furu Wei, · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xintao Wang, Ke Yu, Shixiang Wu, Jinjin Gu, Yihao Liu, Chao Dong, Yu Qiao, and Chen Change Loy, · 2018
Cited alongside, same era.
“A novel two-stage separable deep learning framework for practical blind watermarking,”
Yang Liu, Mengxi Guo, Jian Zhang, Yuesheng Zhu, and Xiaodong Xie, · 2019
Cited alongside, same era.
“Reversible gans for memory-efficient image-to-image translation,”
Tycho FA van der Ouderaa and Daniel E Worrall, · 2019
Cited alongside, same era.
“Waveglow: A flow-based generative network for speech synthesis,”
Ryan Prenger, Rafael Valle, and Bryan Catanzaro, · 2019
Cited alongside, same era.
“Common voice: A massively-multilingual speech corpus,”
Rosana Ardila, Megan Branson, KellyCue Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, LindsayR. Saunders, FrancisM. Tyers, and Gregor Weber, · 2019
Cited alongside, same era.
“Glow-tts: A generative flow for text-to-speech via monotonic alignment search,”
Jaehyeon Kim, Sungwon Kim, Jungil Kong, and Sungroh Yoon, · 2020
Cited alongside, same era.
“Audiowmark: Audio watermarking,” https://uplex.de/audiowmark , 2020
Stefan Westerfeld, · 2020
Cited alongside, same era.
Closest in time.
“Mega-tts: Zero-shot text-to-speech at scale with intrinsic inductive bias,”
Ziyue Jiang, Yi Ren, Zhenhui Ye, Jinglin Liu, Chen Zhang, Qian Yang, Shengpeng Ji, Rongjie Huang, Chunfeng Wang, Xiang Yin, Zejun Ma, and Zhou Zhao, · 2023
Closest in time.
“Make-a-voice: Unified voice synthesis with discrete representation,”
Rongjie Huang, Chunlei Zhang, Yongqi Wang, Dongchao Yang, Luping Liu, Zhenhui Ye, Ziyue Jiang, Chao Weng, Zhou Zhao, and Dong Yu, · 2023
Closest in time.
“Dear: A deep-learning-based audio re-recording resilient watermarking,”
Chang Liu, Jie Zhang, Han Fang, Zehua Ma, Weiming Zhang, and Nenghai Yu, · 2023
Closest in time.
“Simple and controllable music generation,”
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez, · 2023
Closest in time.
“Speak, read and prompt: High-fidelity text-to-speech with minimal supervision,”
Eugene Kharitonov, Damien Vincent, Zalán Borsos, Raphaël Marinier, Sertan Girgin, Olivier Pietquin, Matt Sharifi, Marco Tagliasacchi, and Neil Zeghidour, · 2023
Closest in time.
“Rosteals: Robust steganography using autoencoder latent space,”
Tu Bui, Shruti Agarwal, Ning Yu, and John Collomosse, · 2023
Closest in time.
“Irwart: Levering watermarking performance for protecting high-quality artwork images,”
Yuanjing Luo, Tongqing Zhou, Fang Liu, and Zhiping Cai, · 2023
Closest in time.
“Large-capacity and flexible video steganography via invertible neural network,”
Chong Mou, Youmin Xu, Jiechong Song, Chen Zhao, Bernard Ghanem, and Jian Zhang, · 2023
Closest in time.
“Dnn-based speech watermarking resistant to desynchronization attacks,”
Kosta Pavlović, Slavko Kovačević, Igor Djurović, and Adam Wojciechowski, · 2023
Closest in time.
“Towards robust image-in-audio deep steganography,”
Jaume Ros Alonso, Margarita Geleta, Jordi Pons, and Xavier Giro-i Nieto, · 2023
Closest in time.