Fetching the paper…
Reading the bibliography…
Many audio processing tasks require perceptual assessment.
A. W. Rix and J. G. Beerends et al., “Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,” in
2001
Earlier work this paper cites.
I. Cohen and B. Berdugo, “Speech enhancement for non-stationary noise environments,”
2001
Earlier work this paper cites.
Z. Wang, A. C. Bovik
2004
Earlier work this paper cites.
Y. Hu and P. C. Loizou, “Subjective comparison of speech enhancement algorithms,” in
2006
Earlier work this paper cites.
T. Manjunath, “Limitations of perceptual evaluation of speech quality on voip systems,” in
2009
Earlier work this paper cites.
J. G. Beerends, C. Schmidmer
2013
Earlier work this paper cites.
A. Hines, J. Skoglund, A. Kokaram, and N. Harte, “Robustness of speech quality metrics to background noise and network degradations: Comparing visqol, pesq and polqa,” in
2013
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,”
2014
Earlier work this paper cites.
A. Hines, J. Skoglund, A. C. Kokaram, and N. Harte, “Visqol: an objective speech quality model,”
2015
Earlier work this paper cites.
L. A. Gatys, A. S. Ecker, and M. Bethge, “A neural algorithm of artistic style,”
2015
Earlier work this paper cites.
D. McShefferty, W. M. Whitmer, and M. A. Akeroyd, “The just-noticeable difference in speech-to-noise ratio,”
2015
Earlier work this paper cites.
K. J. Piczak, “Esc: Dataset for environmental sound classification,” in
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
M. Cartwright, B. Pardo, G. J. Mysore, and M. Hoffman, “Fast and easy crowdsourced perceptual audio evaluation,” in
2016
Cited alongside, same era.
J. Traer and J. H. McDermott, “Statistics of natural reverberation enable perceptual separation of sound and space,”
2016
Cited alongside, same era.
C. Valentini-Botinhao, X. Wang, S. Takaki, and J. Yamagishi, “Speech enhancement for a noise-robust text-to-speech synthesis system using deep recurrent neural networks.” in
2016
Cited alongside, same era.
J. Johnson, A. Alahi, and L. Fei-Fei, “Perceptual losses for real-time style transfer and super-resolution,” in
2016
Cited alongside, same era.
S. Pascual, A. Bonafonte, and J. Serra, “Segan: Speech enhancement generative adversarial network,”
2017
Cited alongside, same era.
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in
2018
Later among the works it cites.
F. G. Germain, Q. Chen, and V. Koltun, “Speech denoising with deep feature losses,”
2018
Later among the works it cites.
K. Kilgour, M. Zuluaga, D. Roblek, and M. Sharifi, “Fr
2018
Later among the works it cites.
M. Cartwright, B. Pardo, and G. J. Mysore, “Crowdsourced pairwise-comparison for source separation evaluation,” in
2018
Later among the works it cites.
Z. Jin, A. Finkelstein, G. J. Mysore, and J. Lu, “Fftnet: A real-time speaker-dependent neural vocoder,” in
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Mesaros, T. Heittola, E. Benetos, P. Foster, M. Lagrange, T. Virtanen, and M. D. Plumbley, “Detection and classification of acoustic scenes and events: Outcome of the dcase 2016 challenge,”
2017
Cited alongside, same era.
S. Hershey, S. Chaudhuri, D. P. Ellis
2017
Cited alongside, same era.
J. F. Gemmeke, D. P. Ellis
2017
Cited alongside, same era.
Z. Jin, G. J. Mysore, S. Diverdi, J. Lu, and A. Finkelstein, “Voco: text-based insertion and replacement in audio narration,”
2017
Cited alongside, same era.
H. Zhang, X. Zhang, and G. Gao, “Training supervised speech separation system to improve stoi and pesq directly,” in
2018
Cited alongside, same era.
C. Donahue, J. McAuley, and M. Puckette, “Adversarial audio synthesis,”
2018
Cited alongside, same era.
D. Stoller, S. Ewert, and S. Dixon, “Adversarial semi-supervised audio source separation applied to singing voice extraction,” in
2018
Cited alongside, same era.
D. Rethage, J. Pons, and X. Serra, “A wavenet for speech denoising,” in
2018
Later among the works it cites.
S.-W. Fu, C.-F. Liao, and Y. Tsao, “Learning with learned loss function: Speech enhancement with quality-net to improve perceptual evaluation of speech quality,”
2019
Later among the works it cites.
I. Ananthabhotla, S. Ewert, and J. A. Paradiso, “Towards a perceptual loss: Using a neural network codec approximation as a loss for generative audio models,” in
2019
Later among the works it cites.
A. R. Avila, J. Alam, D. O’Shaughnessy, and T. H. Falk, “Intrusive quality measurement of noisy and enhanced speech based on i-vector similarity,” in
2019
Later among the works it cites.
J. Cramer, H.-H. Wu, J. Salamon, and J. P. Bello, “Look, listen, and learn more: Design choices for deep audio embeddings,” in
2019
Later among the works it cites.
B. Feng, Z. Jin, J. Su, and A. Finkelstein, “Learning bandwidth expansion using perceptually-motivated loss,” in
2019
Later among the works it cites.
H. Purwins, B. Li, T. Virtanen, J. Schlüter, S.-Y. Chang, and T. Sainath, “Deep learning for audio signal processing,”
2019
Later among the works it cites.