Fetching the paper…
Reading the bibliography…
Self-supervised speech model is a rapid progressing research topic, and many pre-trained models have been released and used in various down stream tasks.
“Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences,”
Steven Davis and Paul Mermelstein, · 1980
Earlier work this paper cites.
“SWITCHBOARD: Telephone speech corpus for research and development,”
John J Godfrey, Edward C Holliman, and Jane McDaniel, · 1992
Earlier work this paper cites.
“The Fisher Corpus: a Resource for the Next Generations of Speech-to-Text,”
Christopher Cieri, David Miller, and Kevin Walker, · 2004
Earlier work this paper cites.
“Babel program,”
IARPA, · 2011
Earlier work this paper cites.
“Spoofing and countermeasures for automatic speaker verification,”
Nicholas Evans, Tomi Kinnunen, and Junichi Yamagishi, · 2013
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Librispeech: an ASRcorpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Constant Q cepstral coefficients: A spoofing countermeasure for automatic speaker verification,”
Massimiliano Todisco, Héctor Delgado, and Nicholas Evans, · 2017
Earlier work this paper cites.
“DNN Filter Bank Cepstral Coefficients for Spoofing Detection,”
Hong Yu, Zheng-Hua Tan, Yiming Zhang, Zhanyu Ma, and Jun Guo, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Speaker Recognition from raw waveform with SincNet,”
Mirco Ravanelli and Yoshua Bengio, · 2018
Earlier work this paper cites.
“Long range acoustic and deep features perspective on ASVspoof 2019,”
Rohan Kumar Das, Jichen Yang, and Haizhou Li, · 2019
Earlier work this paper cites.
“ASVspoof 2019: future horizons in spoofed and fake audio detection,”
Massimiliano Todisco, Xin Wang, Ville Vestman, Md. Sahidullah, Héctor Delgado, Andreas Nautsch, Junichi Yamagishi, Nicholas Evans, Tomi H Kinnunen, and Kong Aik Lee, · 2019
Cited alongside, same era.
“fairseq: A Fast, Extensible Toolkit for Sequence Modeling,”
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli, · 2019
Cited alongside, same era.
“SciPy: Open source scientific tools for Python,” 2001
Eric Jones, Travis Oliphant, Pearu Peterson, and Others, · 2019
Cited alongside, same era.
“Advances in anti-spoofing: from the perspective of ASVspoof challenges,”
Madhu R Kamble, Hardik B Sailor, Hemant A Patil, and Haizhou Li, · 2020
Cited alongside, same era.
“End-to-end anti-spoofing with RawNet2,”
Hemlata Tak, Jose Patino, Massimiliano Todisco, Andreas Nautsch, Nicholas Evans, and Anthony Larcher, · 2020
Cited alongside, same era.
“FastAudio: A Learnable Audio Front-End for Spoof Speech Detection,”
Quchen Fu, Zhongwei Teng, Jules White, Maria Powell, and Douglas C Schmidt, · 2021
Closest in time.
“ASVspoof 2021: accelerating progress in spoofed and deepfake speech detection,”
Junichi Yamagishi, Xin Wang, Massimiliano Todisco, Md Sahidullah, Jose Patino, Andreas Nautsch, Xuechen Liu, Kong Aik Lee, Tomi Kinnunen, Nicholas Evans, and Héctor Delgado, · 2021
Closest in time.
“Siamese network with wav2vec feature for spoofing speech detection,”
Yang Xie, Zhenchuan Zhang, and Yingchun Yang, · 2021
Closest in time.
“HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Closest in time.
“SUPERB: Speech Processing Universal PERformance Benchmark,”
Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y Lin, Andy T Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, Tzu-Hsien Huang, Wei-Cheng Tseng, Ko-tik Lee, Da-Rong Liu, Zili Huang, Shuyan Dong, Shang-Wen Li, Shinji Watanabe, Abdelrahman Mohamed, and Hung-yi Lee, · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rohan Kumar Das, Jichen Yang, and Haizhou Li, · 2020
Cited alongside, same era.
“Self-Supervised Spoofing Audio Detection Scheme,”
Ziyue Jiang, Hongcheng Zhu, Li Peng, Wenbing Ding, and Yanzhen Ren, · 2020
Cited alongside, same era.
“wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Cited alongside, same era.
“Common Voice: A Massively-Multilingual Speech Corpus,”
Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis Tyers, and Gregor Weber, · 2020
Cited alongside, same era.
“Libri-light: A benchmark for ASR with limited or no supervision,”
Jacob Kahn, Morgane Rivière, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, Julien Karadayi, Vitaliy Liptchinsky, Ronan Collobert, Christian Fuegen, and Others, · 2020
Cited alongside, same era.
“An Explainability Study of the Constant Q Cepstral Coefficient Spoofing Countermeasure for Automatic Speaker Verification,”
Hemlata Tak, Jose Patino, Andreas Nautsch, Nicholas Evans, and Massimiliano Todisco, · 2020
Cited alongside, same era.
“PyTorch: An Imperative Style, High-Performance Deep Learning Library,”
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala,
Cited in the paper.
Closest in time.
“A comparative study on recent neural spoofing countermeasures for synthetic speech detection,”
Xin Wang and Junich Yamagishi, · 2021
Closest in time.
Jui Shah, Yaman Kumar Singla, Changyou Chen, and Rajiv Ratn Shah, · 2021
Closest in time.
“Layer-wise Analysis of a Self-supervised Speech Representation Model,”
Ankita Pasad, Ju-Chieh Chou, and Karen Livescu, · 2021
Closest in time.
“ASVspoof 2015: the first automatic speaker verification spoofing and countermeasures challenge,”
Zhizheng Wu, Tomi Kinnunen, Nicholas Evans, Junichi Yamagishi, Cemal Hanilçi, Md Sahidullah, and Aleksandr Sizov, · 2041
Closest in time.
“Generalization of spoofing countermeasures: A case study with ASVspoof 2015 and BTAS 2016 corpora,”
Dipjyoti Paul, Md Sahidullah, and Goutam Saha, · 2051
Closest in time.
“A comparison of features for synthetic speech detection,”
Md Sahidullah, Tomi Kinnunen, and Cemal Hanilçi, · 2091
Closest in time.