Fetching the paper…
Reading the bibliography…
In the rapidly evolving field of speech generative models, there is a pressing need to ensure audio authenticity against the risks of voice cloning.
Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs
Rix, A. W., Beerends, J. G., Hollier, M. P., and Hekstra, A. P · 2001
Earlier work this paper cites.
Audio watermark robustness to desynchronization via beat detection
Kirovski, D. and Attias, H · 2003
Earlier work this paper cites.
Spread-spectrum watermarking of audio signals
Kirovski, D. and Malvar, H. S · 2003
Earlier work this paper cites.
Robust and high-quality time-domain audio watermarking based on low-frequency amplitude modification
Lie, W. and Chang, L · 2005
Earlier work this paper cites.
A constructive and unifying framework for zero-bit watermarking
Furon, T · 2007
Earlier work this paper cites.
Robust multiplicative patchwork method for audio watermarking
Kalantari, N. K., Akhaee, M. A., Ahadi, S. M., and Amindavar, H · 2009
Earlier work this paper cites.
A short-time objective intelligibility measure for time-frequency weighted noisy speech
Taal, C. H., Hendriks, R. C., Heusdens, R., and Jensen, J · 2010
Earlier work this paper cites.
Auditory neuroscience: Making sense of sound
Schnupp, J., Nelken, I., and King, A · 2011
Earlier work this paper cites.
Algorithms to measure audio programme loudness and true-peak audio level
telecommunication Union, I · 2011
Earlier work this paper cites.
Visqol: The virtual speech quality objective listener
Hines, A., Skoglund, J., Kokaram, A., and Harte, N · 2012
Earlier work this paper cites.
Robust patchwork-based embedding and decoding scheme for digital audio watermarking
Natgunanathan, I., Xiang, Y., Rong, Y., Zhou, W., and Guo, S · 2012
Earlier work this paper cites.
Method for the subjective assessment of intermediate quality level of audio systems
Series, B · 2014
Earlier work this paper cites.
Patchwork-based audio watermarking method robust to de-synchronization attacks
Xiang, Y., Natgunanathan, I., Guo, S., Zhou, W., and Nahavandi, S · 2014
Earlier work this paper cites.
Spoofing countermeasure based on analysis of linear prediction error
Janicki, A · 2015
Earlier work this paper cites.
A comparison of features for synthetic speech detection
Sahidullah, M., Kinnunen, T., and Hanilçi, C · 2015
Earlier work this paper cites.
Wavenet: A generative model for raw audio
van den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F., Ellis, D. P., Freedman, D., Jansen, A., Lawrence, W., Moore, R. C., Plakal, M., and Ritter, M · 2017
Earlier work this paper cites.
Spread spectrum audio watermarking using multiple orthogonal PN sequences and variable embedding strengths and polarities
Xiang, Y., Natgunanathan, I., Peng, D., Hua, G., and Liu, B · 2017
Earlier work this paper cites.
An investigation of deep-learning frameworks for speaker verification antispoofing
Zhang, C., Yu, C., and Hansen, J. H · 2017
Earlier work this paper cites.
Neural voice cloning with a few samples
Arik, S., Chen, J., Peng, K., Ping, W., and Zhou, Y · 2018
Earlier work this paper cites.
Patchwork-based audio watermarking robust against de-synchronization and recapturing attacks
Liu, Z., Huang, Y., and Huang, J · 2018
Earlier work this paper cites.
Snr-constrained heuristics for optimizing the scaling parameter of robust audio watermarking
Su, Z., Zhang, G., Yue, F., Chang, L., Jiang, J., and Yao, X · 2018
Earlier work this paper cites.
Detecting ai-synthesized speech using bispectral analysis
AlBadawy, E. A., Lyu, S., and Farid, H · 2019
Earlier work this paper cites.
Melgan: Generative adversarial networks for conditional waveform synthesis
Kumar, K., Kumar, R., de Boissière, T., Gestin, L., Teoh, W. Z., Sotelo, J. M. R., de Brébisson, A., Bengio, Y., and Courville, A. C · 2019
Earlier work this paper cites.
Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation
Luo, Y. and Mesgarani, N · 2019
Cited alongside, same era.
Audio watermarking over the air with modulated self-correlation
Tai, Y.-Y. and Mansour, M. F · 2019
Cited alongside, same era.
Understanding straight-through estimator in training activation quantized neural nets
Yin, P., Lyu, J., Zhang, S., Osher, S., Qi, Y., and Xin, J · 2019
Cited alongside, same era.
Real time speech enhancement in the waveform domain, 2020
Defossez, A., Synnaeve, G., and Adi, Y · 2020
Cited alongside, same era.
A spectral energy distance for parallel speech synthesis
Gritsenko, A., Salimans, T., van den Berg, R., Snoek, J., and Kalchbrenner, N · 2020
Cited alongside, same era.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Soundstorm: Efficient parallel audio generation
Borsos, Z., Sharifi, M., Vincent, D., Kharitonov, E., Zeghidour, N., and Tagliasacchi, M · 2023
Later among the works it cites.
Wavmark: Watermarking for audio generation
Chen, G., Wu, Y., Liu, S., Liu, T., Du, X., and Wei, F · 2023
Later among the works it cites.
Simple and controllable music generation
Copet, J., Kreuk, F., Gat, I., Remez, T., Kant, D., Synnaeve, G., Adi, Y., and Défossez, A · 2023
Later among the works it cites.
Three bricks to consolidate watermarks for large language models
Fernandez, P., Chaffin, A., Tit, K., Chappelier, V., and Furon, T · 2023
Later among the works it cites.
Audiobox: Unified audio generation with natural language prompts
Hsu, W.-N., Akinyemi, A., Rakotoarison, A., Tjandra, A., Vyas, A., Guo, B., Akula, B., Shi, B., Ellis, B., Cruz, I., Wang, J., Zhang, J., Williamson, M., Le, M., Moritz, R., Adkins, R., Ngan, W., Zhang, X., Yungster, Y., and Wu, Y.-C · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kong, J., Kim, J., and Bae, J · 2020
Cited alongside, same era.
Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech separation
Luo, Y., Chen, Z., and Yoshioka, T · 2020
Cited alongside, same era.
Synthetic speech detection through short-term and long-term prediction traces
Borrelli, C., Bestagini, P., Antonacci, F., Sarti, A., and Tubaro, S · 2021
Cited alongside, same era.
Fakeavceleb: A novel audio-video multimodal deepfake dataset, 2021
Khalid, H., Tariq, S., and Woo, S. S · 2021
Cited alongside, same era.
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech
Kim, J., Kong, J., and Son, J · 2021
Cited alongside, same era.
Asvspoof 2021: Towards spoofed and deepfake speech detection in the wild
Liu, X., Wang, X., Sahidullah, M., Patino, J., Delgado, H., Kinnunen, T., Todisco, M., Yamagishi, J., Evans, N., Nautsch, A., et al · 2021
Cited alongside, same era.
Voxpopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation
Wang, C., Rivière, M., Lee, A., Wu, A., Talnikar, C., Haziza, D., Williamson, M., Pino, J. M., and Dupoux, E · 2021
Cited alongside, same era.
Later among the works it cites.
Collaborative watermarking for adversarial speech synthesis
Juvela, L. and Wang, X · 2023
Later among the works it cites.
Speak, read and prompt: High-fidelity text-to-speech with minimal supervision
Kharitonov, E., Vincent, D., Borsos, Z., Marinier, R., Girgin, S., Pietquin, O., Sharifi, M., Tagliasacchi, M., and Zeghidour, N · 2023
Later among the works it cites.
Wouaf: Weight modulation for user attribution and fingerprinting in text-to-image diffusion models
Kim, C., Min, K., Patel, M., Cheng, S., and Yang, Y · 2023
Later among the works it cites.
A watermark for large language models
Kirchenbauer, J., Geiping, J., Wen, Y., Katz, J., Miers, I., and Goldstein, T · 2023
Later among the works it cites.
Audiogen: Textually guided audio generation
Kreuk, F., Synnaeve, G., Polyak, A., Singer, U., Défossez, A., Copet, J., Parikh, D., Taigman, Y., and Adi, Y · 2023
Later among the works it cites.
High-fidelity audio compression with improved rvqgan
Kumar, R., Seetharaman, P., Luebs, A., Kumar, I., and Kumar, K · 2023
Later among the works it cites.
Voicebox: Text-guided multilingual universal speech generation at scale
Le, M., Vyas, A., Shi, B., Karrer, B., Sari, L., Moritz, R., Williamson, M., Manohar, V., Adi, Y., Mahadeokar, J., et al · 2023
Later among the works it cites.
Dear: A deep-learning-based audio re-recording resilient watermarking
Liu, C., Zhang, J., Fang, H., Ma, Z., Zhang, W., and Yu, N · 2023
Later among the works it cites.
Expresso: A benchmark and analysis of discrete expressive speech resynthesis
Nguyen, T. A., Hsu, W.-N., d’Avirro, A., Shi, B., Gat, I., Fazel-Zarani, M., Remez, T., Copet, J., Synnaeve, G., Hassid, M., et al · 2023
Later among the works it cites.
Audioqr: Deep neural audio watermarks for qr code
Qu, X., Yin, X., Wei, P., Lu, L., and Ma, Z · 2023
Later among the works it cites.
Who is speaking actually? robust and versatile speaker traceability for voice conversion
Ren, Y., Zhu, H., Zhai, L., Sun, Z., Shen, R., and Wang, L · 2023
Later among the works it cites.
Seamless: Multilingual expressive and streaming speech translation
Seamless Communication, Barrault, L., Chung, Y.-A., Meglioli, M. C., Dale, D., Dong, N., Duppenthaler, M., Duquenne, P.-A., Ellis, B., Elsahar, H., Haaheim, J., Hoffman, J., Hwang, M.-J., Inaguma, H., Klaiber, C., Kulikov, I., Li, P., Licht, D., Maillard, J., Mavlyutov, R., Rakotoarison, A., Sadagopan, K. R., Ramakrishnan, A., Tran, T., Wenzek, G., Yang, Y., Ye, E., Evtimov, I., Fernandez, P., Gao, C., Hansanti, P., Kalbassi, E., Kallet, A., Kozhevnikov, A., Mejia, G., Roman, R. S., Touret, C., Wong, C., Wood, C., Yu, B., Andrews, P., Balioglu, C., Chen, P.-J., Costa-jussà, M. R., Elbayad, M., Gong, H., Guzmán, F., Heffernan, K., Jain, S., Kao, J., Lee, A., Ma, X., Mourachko, A., Peloquin, B., Pino, J., Popuri, S., Ropers, C., Saleem, S., Schwenk, H., Sun, A., Tomasello, P., Wang, C., Wang, J., Wang, S., and Williamson, M · 2023
Later among the works it cites.
Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers
Shen, K., Ju, Z., Tan, X., Liu, Y., Leng, Y., He, L., Qin, T., Zhao, S., and Bian, J · 2023
Later among the works it cites.
Ensuring safe, secure, and trustworthy ai
USA · 2023
Later among the works it cites.
Neural codec language models are zero-shot text to speech synthesizers
Wang, C., Chen, S., Wu, Y., Zhang, Z., Zhou, L., Liu, S., Chen, Z., Liu, Y., Wang, H., Li, J., et al · 2023
Later among the works it cites.
Tree-ring watermarks: Fingerprints for diffusion images that are invisible and robust
Wen, Y., Kirchenbauer, J., Geiping, J., and Goldstein, T · 2023
Later among the works it cites.
Adversarial audio watermarking: Embedding watermark into deep feature
Wu, S., Liu, J., Huang, Y., Guan, H., and Zhang, S · 2023
Later among the works it cites.
Biden audio deepfake spurs ai startup elevenlabs—valued at $1.1 billion—to ban account: ‘we’re going to see a lot more of this’
Murphy, M., Metz, R., Bergen, M., and Bloomberg · 2024
Closest in time.