Fetching the paper…
Reading the bibliography…
Deep generative modeling has the potential to cause significant harm to society.
Communication Theory of Secrecy Systems
Claude E Shannon · 1949
Earlier work this paper cites.
A fast two-dimensional median filtering algorithm
Thomas Huang, GJTGY Yang, and Greory Tang · 1979
Earlier work this paper cites.
Simultaneous Modeling of Spectrum, Pitch and Duration in HMM-Based Speech Synthesis
Takayoshi Yoshimura, Keiichi Tokuda, Takashi Masuko, Takao Kobayashi, and Tadashi Kitamura · 1999
Earlier work this paper cites.
Usability of Biometrics in Relation to Electronic Signatures
Dirk Scheuermann, Scarlet Schwiderski-Grosche, and Bruno Struif · 2000
Earlier work this paper cites.
Speech Parameter Generation Algorithms for HMM-Based Speech Synthesis
Keiichi Tokuda, Takayoshi Yoshimura, Takashi Masuko, Takao Kobayashi, and Tadashi Kitamura · 2000
Earlier work this paper cites.
Advanced Encryption Standard
Vincent Rijmen and Joan Daemen · 2001
Earlier work this paper cites.
The honeynet arms race
Bill McCarty · 2003
Earlier work this paper cites.
Discrete-Time Speech Signal Processing: Principles and Practice
Thomas Quatieri · 2006
Earlier work this paper cites.
Psychoacoustics: Facts and Models
Eberhard Zwicker and Hugo Fastl · 2007
Earlier work this paper cites.
Guide to General Server Security
Karen Scarfone, Wayne Jansen, Miles Tracy, et al · 2008
Earlier work this paper cites.
Statistical Parametric Speech Synthesis
Heiga Zen, Keiichi Tokuda, and Alan W Black · 2009
Earlier work this paper cites.
Outside the Closed World: On Using Machine Learning for Network Intrusion Detection
Robin Sommer and Vern Paxson · 2010
Earlier work this paper cites.
Statistical Parametric Speech Synthesis Using Deep Neural Networks
Heiga Ze, Andrew Senior, and Mike Schuster · 2013
Earlier work this paper cites.
TTS Synthesis with Bidirectional LSTM Based Recurrent Neural Networks
Yuchen Fan, Yao Qian, Feng-Long Xie, and Frank K Soong · 2014
Earlier work this paper cites.
Generative Adversarial Networks
Ian J Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Window Functions and their Applications in Signal Processing
KM Muraleedhara Prabhu · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
A Comparison of Features for Synthetic Speech Detection
Md Sahidullah, Tomi Kinnunen, and Cemal Hanilçi · 2015
Earlier work this paper cites.
Guideline for Using Cryptographic Standards in the Federal Government: Cryptographic Mechanisms
Elaine Barker et al · 2016
Earlier work this paper cites.
Improving Variational Inference with Inverse Autoregressive Flow
Diederik P Kingma, Tim Salimans, Rafal Jozefowicz, Xi Chen, Ilya Sutskever, and Max Welling · 2016
Earlier work this paper cites.
SampleRNN: An Unconditional End-to-End Neural Audio Generation Model
Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron Courville, and Yoshua Bengio · 2016
Earlier work this paper cites.
WaveNet: A Generative Model for Raw Audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Theory and Application of Digital Signal Processing
Lawrence Rabiner, Bernard Gold, and CK Yuen · 2016
Earlier work this paper cites.
Voicelive: A Phoneme Localization Based Liveness Detection for Voice Authentication on Smartphones
Linghan Zhang, Sheng Tan, Jie Yang, and Yingying Chen · 2016
Earlier work this paper cites.
Did you hear that? Adversarial Examples Against Automatic Speech Recognition
Moustafa Alzantot, Bharathan Balaji, and Mani Srivastava · 2017
Earlier work this paper cites.
Deep Voice 2: Multi-Speaker Neural Text-to-Speech
Sercan Arik, Gregory Diamos, Andrew Gibiansky, John Miller, Kainan Peng, Wei Ping, Jonathan Raiman, and Yanqi Zhou · 2017
Earlier work this paper cites.
Deep Voice: Real-Time Neural Text-to-Speech
Sercan Ö Arık, Mike Chrzanowski, Adam Coates, Gregory Diamos, Andrew Gibiansky, Yongguo Kang, Xian Li, John Miller, Andrew Ng, Jonathan Raiman, et al · 2017
Earlier work this paper cites.
The LJ Speech Dataset
Keith Ito and Linda Johnson · 2017
Earlier work this paper cites.
Generating and designing DNA with deep generative models
Nathan Killoran, Leo J Lee, Andrew Delong, David Duvenaud, and Brendan J Frey · 2017
Earlier work this paper cites.
RedDots replayed: A new replay spoofing attack corpus for text-dependent speaker verification research
Tomi Kinnunen, Md Sahidullah, Mauro Falcone, Luca Costantini, Rosa González Hautamäki, Dennis Thomsen, Achintya Sarkar, Zheng-Hua Tan, Héctor Delgado, Massimiliano Todisco, Nicholas Evans, Ville Hautamäki, and Kong Aik Lee · 2017
Earlier work this paper cites.
Audio Replay Attack Detection with Deep Learning Frameworks
Galina Lavrentyeva, Sergey Novoselov, Egor Malykh, Alexander Kozlov, Oleg Kudashev, and Vadim Shchemelinin · 2017
Earlier work this paper cites.
Neural Discrete Representation Learning
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
Novel Variable Length Teager Energy Separation Based Instantaneous Frequency Features for Replay Detection
Hemant A Patil, Madhu R Kamble, Tanvina B Patel, and Meet H Soni · 2017
Earlier work this paper cites.
Spoofing Detection via Simultaneous Verification of Audio-Visual Synchronicity and Transcription
Lea Schönherr, Steffen Zeiler, and Dorothea Kolossa · 2017
Earlier work this paper cites.
JSUT Corpus: Free Large-Scale Japanese Speech Corpus for End-to-End Speech Synthesis
Ryosuke Sonobe, Shinnosuke Takamichi, and Hiroshi Saruwatari · 2017
Earlier work this paper cites.
Char2wav: End-to-End Speech Synthesis
Jose Sotelo, Soroush Mehri, Kundan Kumar, Joao Felipe Santos, Kyle Kastner, Aaron Courville, and Yoshua Bengio · 2017
Earlier work this paper cites.
VoiceLoop: Voice Fitting and Synthesis via a Phonological Loop
Yaniv Taigman, Lior Wolf, Adam Polyak, and Eliya Nachmani · 2017
Earlier work this paper cites.
ASVspoof: The Automatic Speaker Verification Spoofing and Countermeasures Challenge
Zhizheng Wu, Junichi Yamagishi, Tomi Kinnunen, Cemal Hanilçi, Mohammed Sahidullah, Aleksandr Sizov, Nicholas Evans, Massimiliano Todisco, and Héctor Delgado · 2017
Earlier work this paper cites.
Audio Adversarial Examples: Targeted Attacks on Speech-to-Text
Nicholas Carlini and David Wagner · 2018
Cited alongside, same era.
Real-Valued (Medical) Time Series Generation with Recurrent Conditional GANs
Cristóbal Esteban, Stephanie L Hyland, and Gunnar Rätsch · 2018
Cited alongside, same era.
GAN-Based synthetic Medical Image Augmentation for Increased CNN Performance in Liver Lesion Classification
Maayan Frid-Adar, Idit Diamant, Eyal Klang, Michal Amitai, Jacob Goldberger, and Hayit Greenspan · 2018
Cited alongside, same era.
Efficient Neural Audio Synthesis
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron Oord, Sander Dieleman, and Koray Kavukcuoglu · 2018
Cited alongside, same era.
Effectiveness of Speech Demodulation-Based Features for Replay Detection
Madhu R Kamble, Hemlata Tak, and Hemant A Patil · 2018
Cited alongside, same era.
SoK: The Faults in our ASRs: An Overview of Attacks against Automatic Speech Recognition and Speaker Identification Systems
Hadi Abdullah, Kevin Warren, Vincent Bindschaedler, Nicolas Papernot, and Patrick Traynor · 2020
Later among the works it cites.
VENOMAVE: Clean-Label Poisoning Against Speech Recognition
Hojjat Aghakhani, Thorsten Eisenhofer, Lea Schönherr, Dorothea Kolossa, Thorsten Holz, Christopher Kruegel, and Giovanni Vigna · 2020
Later among the works it cites.
Void: A fast and light voice liveness detection system
Muhammad Ejaz Ahmed, Il-Youp Kwak, Jun Ho Huh, Iljoo Kim, Taekkyung Oh, and Hyoungshick Kim · 2020
Later among the works it cites.
A Neural Vocoder With Hierarchical Generation of Amplitude and Phase Spectra for Statistical Parametric Speech Synthesis
Yang Ai and Zhen-Hua Ling · 2020
Later among the works it cites.
Common Voice: A Massively-Multilingual Speech Corpus
Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M Tyers, and Gregor Weber · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuezun Li and Siwei Lyu · 2018
Cited alongside, same era.
Meltdown: Reading kernel memory from user space
Moritz Lipp, Michael Schwarz, Daniel Gruss, Thomas Prescher, Werner Haas, Anders Fogh, Jann Horn, Stefan Mangard, Paul Kocher, Daniel Genkin, Yuval Yarom, and Mike Hamburg · 2018
Cited alongside, same era.
Detection of GAN-Generated Fake Images over Social Networks
Francesco Marra, Diego Gragnaniello, Davide Cozzolino, and Luisa Verdoliva · 2018
Cited alongside, same era.
Detecting GAN-Generated Imagery Using Color Cues
Scott McCloskey and Michael Albright · 2018
Cited alongside, same era.
Fake Faces Identification via Convolutional Neural Network
Huaxiao Mo, Bolin Chen, and Weiqi Luo · 2018
Cited alongside, same era.
Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al · 2018
Cited alongside, same era.
End-To-End Audio Replay Attack Detection Using Deep Convolutional Networks with Attention
Francis Tom, Mohit Jain, and Prasenjit Dey · 2018
Cited alongside, same era.
Later among the works it cites.
High Fidelity Speech Synthesis with Adversarial Networks
Mikołaj Bińkowski, Jeff Donahue, Sander Dieleman, Aidan Clark, Erich Elsen, Norman Casagrande, Luis C Cobo, and Karen Simonyan · 2020
Later among the works it cites.
Telegram Still Hasn’t Removed an AI Bot That’s Abusing Women
Matt Burgess · 2020
Later among the works it cites.
Evading DeepFake-Image Detectors with White-and Black-Box Attacks
Nicholas Carlini and Hany Farid · 2020
Later among the works it cites.
WaveGrad: Estimating Gradients for Waveform Generation
Nanxin Chen, Yu Zhang, Heiga Zen, Ron J Weiss, Mohammad Norouzi, and William Chan · 2020
Later among the works it cites.
Recurrent Convolutional Structures for Audio Spoof and Video DeepFake Detection
Akash Chintha, Bao Thai, Saniat Javid Sohrawardi, Kartavya Bhatt, Andrea Hickerson, Matthew Wright, and Raymond Ptucha · 2020
Later among the works it cites.
The DeepFake Detection Challenge (DFDC) Dataset, 2020
Brian Dolhansky, Joanna Bitton, Ben Pflaum, Jikuo Lu, Russ Howes, Menglin Wang, and Cristian Canton Ferrer · 2020
Later among the works it cites.
Watch your Up-Convolution: CNN Based Generative Deep Neural Networks are Failing to Reproduce Spectral Distributions
Ricard Durall, Margret Keuper, and Janis Keuper · 2020
Later among the works it cites.
Listen to This Deepfake Audio Impersonating a CEO in Brazen Fraud Attempt
Lorenzo Franceschi-Bicchierai · 2020
Later among the works it cites.
Leveraging Frequency Analysis for Deep Fake Image Recognition
Joel Frank, Thorsten Eisenhofer, Lea Schönherr, Asja Fischer, Dorothea Kolossa, and Thorsten Holz · 2020
Later among the works it cites.
Conformer: Convolution-Augmented Transformer for Speech Recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al · 2020
Later among the works it cites.
Parallel WaveGAN (+ MelGAN & Multi-band MelGAN) implementation with Pytorch
Tomoki Hayashi · 2020
Later among the works it cites.
Espnet-TTS: Unified, reproducible, and integratable open source end-to-end text-to-speech toolkit
Tomoki Hayashi, Ryuichi Yamamoto, Katsuki Inoue, Takenori Yoshimura, Shinji Watanabe, Tomoki Toda, Kazuya Takeda, Yu Zhang, and Xu Tan · 2020
Later among the works it cites.
ESPnet-ST: All-in-one speech translation toolkit
Hirofumi Inaguma, Shun Kiyono, Kevin Duh, Shigeki Karita, Nelson Yalta, Tomoki Hayashi, and Shinji Watanabe · 2020
Later among the works it cites.
Improved RawNet with Feature Map Scaling for Text-independent Speaker Verification using Raw Waveforms
Jee-weon Jung, Seung-bin Kim, Hye-jin Shim, Ju-ho Kim, and Ha-Jin Yu · 2020
Later among the works it cites.
Celeb-DF: A New Dataset for DeepFake Forensics
Yuezun Li, Xin Yang, Pu Sun, Honggang Qi, and Siwei Lyu · 2020
Later among the works it cites.
Non-Autoregressive Neural Text-to-Speech
Kainan Peng, Wei Ping, Zhao Song, and Kexin Zhao · 2020
Later among the works it cites.
WaveFlow: A Compact Flow-based Model for Raw Audio
Wei Ping, Kainan Peng, Kexin Zhao, and Zhao Song · 2020
Later among the works it cites.
Thinking in Frequency: Face Forgery Detection by Mining Frequency-Aware Clues
Yuyang Qian, Guojun Yin, Lu Sheng, Zixuan Chen, and Jing Shao · 2020
Later among the works it cites.
FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2020
Later among the works it cites.
Imperio: Robust Over-the-Air Adversarial Examples for Automatic Speech Recognition Systems
Lea Schönherr, Thorsten Eisenhofer, Steffen Zeiler, Thorsten Holz, and Dorothea Kolossa · 2020
Later among the works it cites.
CNN-generated images are surprisingly easy to spot… for now
Sheng-Yu Wang, Oliver Wang, Richard Zhang, Andrew Owens, and Alexei A Efros · 2020
Later among the works it cites.
Attribution in Scale and Space
Shawn Xu, Subhashini Venugopalan, and Mukund Sundararajan · 2020
Later among the works it cites.
Parallel WaveGAN: A Fast Waveform Generation Model Based on Generative Adversarial Networks with Multi-Resolution Spectrogram
Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim · 2020
Later among the works it cites.
End-to-End Adversarial Text-to-Speech
Jeff Donahue, Sander Dieleman, Mikołaj Bińkowski, Erich Elsen, and Karen Simonyan · 2021
Closest in time.
DiffWave: A Versatile Diffusion Model for Audio Synthesis
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro · 2021
Closest in time.
Inauthentic Instagram accounts with synthetic faces target Navalny protests
The Atlantic Council’s Digital Forensic Research Lab · 2021
Closest in time.
ESPnet-SE: End-to-end speech enhancement and separation toolkit designed for ASR integration
Chenda Li, Jing Shi, Wangyou Zhang, Aswin Shanmugam Subramanian, Xuankai Chang, Naoyuki Kamo, Moto Hira, Tomoki Hayashi, Christoph Boeddeker, Zhuo Chen, and Shinji Watanabe · 2021
Closest in time.
Tigray conflict: The fake UN diplomat and other misleading stories
Peter Mwai · 2021
Closest in time.
ASVspoof 2019: Spoofing Countermeasures for the Detection of Synthesized, Converted and Replayed Speech
Andreas Nautsch, Xin Wang, Nicholas Evans, Tomi H. Kinnunen, Ville Vestman, Massimiliano Todisco, Héctor Delgado, Md Sahidullah, Junichi Yamagishi, and Kong Aik Lee · 2021
Closest in time.
End-to-End anti-spoofing with RawNet2
Hemlata Tak, Jose Patino, Massimiliano Todisco, Andreas Nautsch, Nicholas Evans, and Anthony Larcher · 2021
Closest in time.
Multi-Band Melgan: Faster Waveform Generation For High-Quality Text-To-Speech
Geng Yang, Shan Yang, Kai Liu, Peng Fang, Wei Chen, and Lei Xie · 2021
Closest in time.
Project Zero Policy and Disclosure: 2020 Edition, 2020
Tim Willis · 2026
Closest in time.