Fetching the paper…
Reading the bibliography…
Automatic speaker verification (ASV) is one of the most natural and convenient means of biometric person recognition.
K. Tanaka, H. Kameoka, T. Kaneko, N. Hojo, Wavecyclegan2: Time-domain neural post-filter for speech waveform generation, CoRR abs/1904.02892 · 1904
Earlier work this paper cites.
W.-C. Huang, Y.-C. Wu, K. Kobayashi, Y.-H. Peng, H.-T. Hwang, P. Lumban Tobing, Y. Tsao, H.-M. Wang, T. Toda, Generalization of spectrum differential based direct waveform modification for voice conversion, in: Proc. SSW10, 2019 · 1907
Earlier work this paper cites.
J. B. Allen, D. A. Berkley, Image Method for Efficiently Simulating Small-Room Acoustics, J. Acoust. Soc. Am 65 (4) (1979) 943–950
1979
Earlier work this paper cites.
D. Griffin, J. Lim, Signal estimation from modified short-time Fourier transform, IEEE Trans. ASSP 32 (2) (1984) 236–243
1984
Earlier work this paper cites.
D. W. Griffin, J. S. Lim, Signal estimation from modified short-time Fourier transform, IEEE Transactions on Acoustics, Speech, and Signal Processing 32 (2) (1984) 236–243
1984
Earlier work this paper cites.
L.-J. Liu, Z.-H. Ling, Y. Jiang, M. Zhou, L.-R. Dai, WaveNet vocoder with limited training data for voice conversion, in: Annual Conference of the International Speech Communication Association, 2018, pp. 1983–1987
1987
Earlier work this paper cites.
W. Verhelst, M. Roelands, An overlap-add technique based on waveform similarity (wsola) for high quality time-scale modification of speech, in: 1993 IEEE International Conference on Acoustics, Speech, and Signal Processing, Vol. 2, IEEE, 1993, pp. 554–557
1993
Earlier work this paper cites.
C. Veaux, J. Yamagishi, K. MacDonald, CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkitHttp://dx.doi.org/10.7488/ds/1994
1994
Earlier work this paper cites.
K. Tokuda, T. Kobayashi, T. Masuko, S. Imai, Mel-generalized cepstral analysis-a unified approach to speech spectral estimation, in: Third International Conference on Spoken Language Processing, 1994
1994
Earlier work this paper cites.
B. L. Pellom, J. H. Hansen, An experimental study of speaker verification sensitivity to computer voice-altered imposters, in: 1999 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings. ICASSP99 (Cat. No. 99CH36258), Vol. 2, IEEE, 1999, pp. 837–840
1999
Earlier work this paper cites.
H. Kawahara, I. Masuda-Katsuse, A. D. Cheveigné, Restructuring speech representations using a pitch-adaptive time-frequency smoothing and an instantaneous-frequency based F0 extraction: Possible role of a repetitive structure in sounds, Speech Communication 27 (3–4) (1999) 187–207
1999
Earlier work this paper cites.
doi:10.1109/ICASSP.2006.1660175
D. Matrouf, J. . Bonastre, C. Fredouille, Effect of speech transformation on impostor acceptance, in: 2006 IEEE International Conference on Acoustics Speech and Signal Processing Proceedings, Vol. 1, 2006, pp. I–I · 2006
Earlier work this paper cites.
A. Hatch, S. Kajarekar, A. Stolcke, Within-class covariance normalization for svm-based speaker recognition, Vol. 3, 2006
2006
Earlier work this paper cites.
S. Ioffe, Probabilistic linear discriminant analysis, in: European Conference on Computer Vision, Springer, 2006, pp. 531–542
2006
Earlier work this paper cites.
S. J. D. Prince, J. H. Elder, Probabilistic linear discriminant analysis for inferences about identity, in: IEEE 11th International Conference on Computer Vision, ICCV 2007, Rio de Janeiro, Brazil, October 14-20, 2007, 2007, pp. 1–8
2007
Earlier work this paper cites.
S. J. Prince, J. H. Elder, Probabilistic linear discriminant analysis for inferences about identity, in: 2007 IEEE 11th International Conference on Computer Vision, IEEE, 2007, pp. 1–8
2007
Earlier work this paper cites.
A. Graves, Supervised Sequence Labelling with Recurrent Neural Networks, Ph.D. thesis, Technische Universität München (2008)
2008
Earlier work this paper cites.
L. van der Maaten, G. Hinton, Visualizing data using t-{SNE}, Journal of Machine Learning Research 9 (Nov) (2008) 2579–2605
2008
Earlier work this paper cites.
F. Toole, Sound Reproduction: Loudspeakers and Rooms , Audio Engineering Society Presents Series, Elsevier, 2008. URL https://books.google.fr/books?id=sGmz0yONYFcC
2008
Earlier work this paper cites.
E. Vincent, Roomsimove (2008). URL http://homepages.loria.fr/evincent/software/Roomsimove_1.4.zip
2008
Earlier work this paper cites.
P. Kenny, Bayesian speaker verification with heavy-tailed priors, in: Odyssey 2010: The Speaker and Language Recognition Workshop, Brno, Czech Republic, June 28 - July 1, 2010, 2010, p. 14
2010
Earlier work this paper cites.
M. Schröder, M. Charfuelan, S. Pammi, I. Steiner, Open source voice creation toolkit for the MARY TTS platform , in: Interspeech, Florence, Italy, 2011, pp. 3253–3256. URL http://www.isca-speech.org/archive/interspeech_2011/i11_3253.html
2011
Cited alongside, same era.
N. Dehak, P. Kenny, R. Dehak, P. Dumouchel, P. Ouellet, Front-end factor analysis for speaker verification 19 (4) (2011) 788–798
2011
Cited alongside, same era.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, et al., The Kaldi speech recognition toolkit, Tech. rep., IEEE Signal Processing Society (2011)
2011
Cited alongside, same era.
P. Kenny, A small footprint i-vector extractor, in: Proc. Odyssey 2012: the Speaker and Language Recognition Workshop, Singapore, 2012
2012
Cited alongside, same era.
Z. Wu, J. Yamagishi, T. Kinnunen, C. Hanilçi, M. Sahidullah, A. Sizov, N. Evans, M. Todisco, H. Delgado, ASVspoof: the automatic speaker verification spoofing and countermeasures challenge, IEEE Journal of Selected Topics in Signal Processing 11 (4) (2017) 588–604
2017
Later among the works it cites.
T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, S. Khudanpur, A study on data augmentation of reverberant speech for robust speech recognition, in: Proc. IEEE Int. Conf. on Acoustics, Speech and Signal Processing (ICASSP), 2017, pp. 5220–5224
2017
Later among the works it cites.
M. Todisco, H. Delgado, N. Evans, Constant Q cepstral coefficients: A spoofing countermeasure for automatic speaker verification, Computer Speech & Language 45 (2017) 516 – 535
2017
Later among the works it cites.
A. Rosenberg, B. Ramabhadran, Bias and statistical significance in evaluating speech synthesis with Mean Opinion Scores, in: Prof. Interspeech, 2017, pp. 3976–3980
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Evans, T. Kinnunen, J. Yamagishi, Spoofing and countermeasures for automatic speaker verification, in: Proc. Interspeech, Annual Conf. of the Int. Speech Comm. Assoc., Lyon, France, 2013, pp. 925–929
2013
Cited alongside, same era.
H. Zen, A. Senior, M. Schuster, Statistical parametric speech synthesis using deep neural networks, in: Proc. ICASSP, 2013, pp. 7962–7966
2013
Cited alongside, same era.
HTS Working Group, The English TTS system Flite+HTS_engine (2014). URL http://hts-engine.sourceforge.net/
2014
Cited alongside, same era.
HTS Working Group, An example of context-dependent label format for HMM-based speech synthesis in Japanese (2015). URL http://hts.sp.nitech.ac.jp/archives/2.3/HTS-demo_NIT-ATR503-M001.tar.bz2
2015
Cited alongside, same era.
Y. Agiomyrgiannakis, Vocaine the vocoder and applications in speech synthesis, in: 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2015, pp. 4230–4234
2015
Cited alongside, same era.
Y. Li, K. Swersky, R. Zemel, Generative moment matching networks, in: International Conference on Machine Learning, 2015, pp. 1718–1727
2015
Cited alongside, same era.
ISO/IEC 30107. Information technology – biometric presentation attack detection, Standard (2016)
2016
Cited alongside, same era.
M. Morise, F. Yokomori, K. Ozawa, WORLD: A vocoder-based high-quality speech synthesis system for real-time applications, IEICE Trans. on Information and Systems 99 (7) (2016) 1877–1884
2016
Cited alongside, same era.
Md Sahidullah, Hector Delgado, Massimiliano Todisco, Tomi Kinnunen, Nicholas Evans, Junichi Yamagishi, Kong-Aik Lee, Introduction to voice presentation attack detection and recent advances, Book chapter N°15 of "Handbook of Biometric Anti-Spoofing: Presentation Attack Detection"; Springer Marcel, S., Nixon, M.S., Fierrez, J., Evans, N. (Eds.); Springer, 2018, 2018
2018
Later among the works it cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, et al., Natural tts synthesis by conditioning wavenet on mel spectrogram predictions, in: 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2018, pp. 4779–4783
2018
Later among the works it cites.
T. Kinnunen, K. Lee, H. Delgado, N. Evans, M. Todisco, M. Sahidullah, J. Yamagishi, D. A. Reynolds, t-DCF: a detection cost function for the tandem assessment of spoofing countermeasures and automatic speaker verification, in: Proc. Odyssey, Les Sables d’Olonne, France, 2018
2018
Later among the works it cites.
X. Wang, J. Lorenzo-Trueba, S. Takaki, L. Juvela, J. Yamagishi, A comparison of recent waveform generation and acoustic modeling methods for neural-network-based speech synthesis, in: Proc. ICASSP, 2018, pp. 4804–4808
2018
Later among the works it cites.
I. Steiner, S. Le Maguer, Creating new language and voice components for the updated MaryTTS text-to-speech synthesis platform , in: 11th Language Resources and Evaluation Conference (LREC), Miyazaki, Japan, 2018, pp. 3171–3175. URL http://www.lrec-conf.org/proceedings/lrec2018/summaries/1045.html
2018
Later among the works it cites.
W.-C. Huang, H.-T. Hwang, Y.-H. Peng, Y. Tsao, H.-M. Wang, Voice conversion based on cross-domain features using variational auto encoders, in: 2018 11th International Symposium on Chinese Spoken Language Processing (ISCSLP), IEEE, 2018, pp. 51–55
2018
Later among the works it cites.
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, P. Nguyen, R. Pang, I. L. Moreno, Y. Wu, et al., Transfer learning from speaker verification to multispeaker text-to-speech synthesis, in: Advances in Neural Information Processing Systems, 2018, pp. 4480–4490
2018
Later among the works it cites.
L. Wan, Q. Wang, A. Papir, I. L. Moreno, Generalized end-to-end loss for speaker verification, in: 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2018, pp. 4879–4883
2018
Later among the works it cites.
K. Kobayashi, T. Toda, S. Nakamura, Intra-gender statistical singing voice conversion with direct waveform modification using log-spectral differential, Speech Communication 99 (2018) 211–220
2018
Later among the works it cites.
doi:10.21437/Odyssey.2018-27
T. Kinnunen, J. Lorenzo-Trueba, J. Yamagishi, T. Toda, D. Saito, F. Villavicencio, Z. Ling, A spoofing benchmark for the 2018 voice conversion challenge: Leveraging from spoofing countermeasures for speech artifact assessment , in: Proc. Odyssey 2018 The Speaker and Language Recognition Workshop, 2018, pp. 187–194 · 2018
Later among the works it cites.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, S. Khudanpur, X-vectors: Robust DNN embeddings for speaker recognition, in: Proc. IEEE Int. Conf. on Acoustics, Speech and Signal Processing (ICASSP), 2018, pp. 5329–5333
2018
Later among the works it cites.
A. organization team, ASVspoof 2019: the automatic speaker verification spoofing and countermeasures challenge evaluation plan . URL http://www.asvspoof.org/asvspoof2019/asvspoof2019_evaluation_plan.pdf
2019
Closest in time.
doi:10.21437/Interspeech.2019-2249
M. Todisco, X. Wang, V. Vestman, M. Sahidullah, H. Delgado, A. Nautsch, J. Yamagishi, N. Evans, T. H. Kinnunen, K. A. Lee, ASVspoof 2019: future horizons in spoofed and fake audio detection , in: Proc. Interspeech, 2019, pp. 1008–1012 · 2019
Closest in time.
doi:10.1109/ICASSP.2019.8682298
X. Wang, S. Takaki, J. Yamagishi, Neural source-filter-based waveform model for statistical parametric speech synthesis, in: ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019, pp. 5916–5920 · 2019
Closest in time.
P.-M. Bousquet, M. Rouvier, On robustness of unsupervised domain adaptation for speaker recognition, 2019, pp. 2958–2962
2019
Closest in time.
Z. Wu, T. Kinnunen, N. Evans, J. Yamagishi, C. Hanilçi, M. Sahidullah, A. Sizov, ASVspoof 2015: the first automatic speaker verification spoofing and countermeasures challenge, in: Proc. Interspeech, Annual Conf. of the Int. Speech Comm. Assoc., Dresden, Germany, 2015, pp. 2037–2041
2041
Closest in time.
M. Sahidullah, T. Kinnunen, C. Hanilçi, A comparison of features for synthetic speech detection, in: Proc. Interspeech, Annual Conf. of the Int. Speech Comm. Assoc., Dresden, Germany, 2015, pp. 2087–2091
2091
Closest in time.