Fetching the paper…
Reading the bibliography…
Many speech processing methods based on deep learning require an automatic and differentiable audio metric for the loss function.
“Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment,”
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, · 2001
Earlier work this paper cites.
“Subjective comparison and evaluation of speech enhancement algorithms,”
Y. Hu and P. C. Loizou, · 2007
Earlier work this paper cites.
“Subjective and objective quality assessment of audio source separation,”
V. Emiya, E. Vincent, N. Harlander, et al., · 2011
Earlier work this paper cites.
“Robustness of speech quality metrics: Comparing ViSQOL, PESQ and POLQA,”
A. Hines, J. Skoglund, A. Kokaram, and N. Harte, · 2013
Earlier work this paper cites.
“Representation learning: A review and new perspectives,”
Y. Bengio, A. Courville, and P. Vincent, · 2013
Earlier work this paper cites.
“Can we automatically transform speech recorded on common consumer devices in real-world environments into professional production quality speech?—a dataset, insights, and challenges,”
G J. Mysore, · 2014
Earlier work this paper cites.
“ViSQOL: an objective speech quality model,”
A. Hines, J. Skoglund, A C. Kokaram, et al., · 2015
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
D P. Kingma and J. Ba, · 2015
Earlier work this paper cites.
“AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech,”
B. Patton, Y. Agiomyrgiannakis, M. Terry, K. Wilson, R. A. Saurous, and D. Sculley, · 2016
Earlier work this paper cites.
“Voco: Text-based insertion and replacement in audio narration,”
Z. Jin, G J. Mysore, S. Diverdi, J. Lu, and A. Finkelstein, · 2017
Earlier work this paper cites.
“The LJ speech dataset,” 2017
K. Ito et al., · 2017
Earlier work this paper cites.
“Noisy speech database for training speech enhancement algorithms and TTS models,”
C. Valentini-Botinhao et al., · 2017
Earlier work this paper cites.
“Towards a definition of disentangled representations,”
I. Higgins, D. Amos, D. Pfau, S. Racaniere, L. Matthey, D. Rezende, and A. Lerchner, · 2018
Earlier work this paper cites.
“Training supervised speech separation system to improve STOI and PESQ directly,”
H. Zhang, X. Zhang, and G. Gao, · 2018
Cited alongside, same era.
“Content-based representations of audio using siamese neural networks,”
P. Manocha, R. Badlani, A. Kumar, A. Shah, B. Elizalde, and B. Raj, · 2018
Cited alongside, same era.
“Learning disentangled representations for timber and pitch in music audio,”
Y N. Hung, Y A. Chen, and Y H. Yang, · 2018
Cited alongside, same era.
J. Chou, C. Yeh, H. Lee, and L. Lee, · 2018
Cited alongside, same era.
“Representation learning with contrastive predictive coding,”
A. Oord, Y. Li, and O. Vinyals, · 2018
“wav2vec: Unsupervised pre-training for speech recognition,”
S. Schneider, A. Baevski, R. Collobert, and M. Auli, · 2019
Later among the works it cites.
“Learning bandwidth expansion using perceptually-motivated loss,”
B. Feng, Z. Jin, J. Su, and A. Finkelstein, · 2019
Later among the works it cites.
“Perceptually-motivated environment-specific speech enhancement,”
J. Su, A. Finkelstein, and Z. Jin, · 2019
Later among the works it cites.
“A differentiable perceptual audio metric learned from just noticeable differences,”
P. Manocha, A. Finkelstein, et al., · 2020
Later among the works it cites.
“SESQA: semi-supervised learning for speech quality assessment,”
J. Serrà, J. Pons, et al., · 2020
Later among the works it cites.
“Real time speech enhancement in the waveform domain,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“FFTNet: A real-time speaker-dependent neural vocoder,”
Z. Jin, A. Finkelstein, G J. Mysore, and J. Lu, · 2018
Cited alongside, same era.
“The voice conversion challenge 2018,”
J. Lorenzo-Trueba, J. Yamagishi, T. Toda, D. Saito, F. Villavicencio, T. Kinnunen, and Z. Ling, · 2018
Cited alongside, same era.
“MOSNet: Deep learning based objective assessment for voice conversion,”
C-C. Lo, S-W. Fu, W-C. Huang, X. Wang, Junichi Yamagishi, Yu Tsao, and Hsin-Min Wang, · 2019
Cited alongside, same era.
“Learning with learned loss function: Speech enhancement with quality-net,”
S-W. Fu, C-F. Liao, and Y. Tsao, · 2019
Cited alongside, same era.
“MetricGAN: Generative adversarial networks based black-box metric scores optimization for speech enhancement,”
S-W. Fu, C-F. Liao, Y. Tsao, and S.D. Lin, · 2019
Cited alongside, same era.
“Melgan: Generative adversarial networks for conditional waveform synthesis,”
K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. de Brebisson, Y. Bengio, and A. Courville, · 2019
Cited alongside, same era.
“Are disentangled representations helpful for abstract visual reasoning?,”
S V. Steenkiste, F. Locatello, J. Schmidhuber, and O. Bachem, · 2019
Cited alongside, same era.
A. Defossez, G. Synnaeve, and Y. Adi, · 2020
Later among the works it cites.
“Unsupervised cross-lingual representation learning for speech recognition,”
A. Conneau, A. Baevski, et al., · 2020
Later among the works it cites.
“Disentangled multidimensional metric learning for music similarity,”
J. Lee, N J. Bryan, J. Salamon, Z. Jin, and J. Nam, · 2020
Later among the works it cites.
“Audio Albert: A lite BERT for self-supervised learning of audio representation,”
P. Chi, P. Chung, T. Wu, et al., · 2020
Later among the works it cites.
“Self-supervised contrastive learning for unsupervised phoneme segmentation,”
F. Kreuk, J. Keshet, and Y. Adi, · 2020
Later among the works it cites.
“A simple framework for contrastive learning of visual representations,”
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, · 2020
Later among the works it cites.
“HiFi-GAN: High-fidelity denoising and dereverberation,”
J. Su, Z. Jin, et al., · 2020
Later among the works it cites.
“The interspeech 2020 deep noise suppression challenge: Datasets, subjective speech quality and testing framework,”
C K. Reddy, E. Beyrami, H. Dubey, V. Gopal, et al., · 2020
Later among the works it cites.