Fetching the paper…
Reading the bibliography…
We present a method to separate speech signals from noisy environments in the embedding space of a neural audio codec.
S. Morishima, H. Harashima, and Y. Katayama, “Speech coding based on a multi-layer neural network,” in IEEE International Conference on Communications , vol. 2, 1990, pp. 429–433
1990
Earlier work this paper cites.
P. C. Loizou, Speech Enhancement: Theory and Practice . Boca Raton: CRC Press, 2007
2007
Earlier work this paper cites.
J.-M. Valin, K. Vos, and T. Terriberry, “Definition of the Opus audio codec,” Internet Engineering Task Force, Request for Comments RFC 6716, 2012
2012
Earlier work this paper cites.
J. Kearns, “LibriVox: Free public domain audiobooks,” Reference Reviews , vol. 28, pp. 7–8, 2014
2014
Earlier work this paper cites.
T. D. Kulkarni, W. Whitney, P. Kohli, and J. B. Tenenbaum, “Deep convolutional inverse graphics network,” in Advances in Neural Information Processing Systems , vol. 28, 2015
2015
Earlier work this paper cites.
A. Hines, J. Skoglund, A. C. Kokaram, and N. Harte, “ViSQOL: An objective speech quality model,” EURASIP Journal on Audio, Speech, and Music Processing , p. 13, 2015
2015
Earlier work this paper cites.
G. Lample, N. Zeghidour, N. Usunier, A. Bordes, L. Denoyer, and M. Ranzato, “Fader networks: Manipulating images by sliding attributes,” in Advances in Neural Information Processing Systems , vol. 30, 2017
2017
Earlier work this paper cites.
E. Fonseca, J. Pons Puig, X. Favory, F. Font Corbera, D. Bogdanov, A. Ferraro, S. Oramas, A. Porter, and X. Serra, “Freesound datasets: A platform for the creation of open audio datasets,” in International Society for Music Information Retrieval Conference , 2017, pp. 486–493
2017
Earlier work this paper cites.
S. Kankanahalli, “End-To-end optimized speech coding with deep neural networks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 2521–2525
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Q. Hu, A. Szabó, T. Portenier, M. Zwicker, and P. Favaro, “Disentangling factors of variation by mixing them,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018, pp. 3399–3407
2018
Cited alongside, same era.
C. Gârbacea, A. van den Oord, Y. Li, F. S. C. Lim, A. Luebs, O. Vinyals, and T. C. Walters, “Low bit-rate speech coding with VQ-VAE and a WaveNet decoder,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 735–739
2019
Cited alongside, same era.
S. Bhagat, V. Udandarao, S. Uppal, and S. Anand, “DisCont: Self-supervised visual attribute disentanglement using context vectors,” in European Conference on Computer Vision Workshops , A. Bartoli and A. Fusiello, Eds., Cham, 2020, pp. 549–553
2020
Later among the works it cites.
B. Gfeller, C. Frank, D. Roblek, M. Sharifi, M. Tagliasacchi, and M. Velimirović, “SPICE: Self-supervised pitch estimation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 1118–1128, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
A. A. Gritsenko, T. Salimans, R. van den Berg, J. Snoek, and N. Kalchbrenner, “A spectral energy distance for parallel speech synthesis,” in Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 13 062–13 072
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
I. Kavalerov, S. Wisdom, H. Erdogan, B. Patton, K. Wilson, J. Le Roux, and J. R. Hershey, “Universal sound separation,” in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) , 2019, pp. 175–179
2019
Cited alongside, same era.
K. Qian, Y. Zhang, S. Chang, X. Yang, and M. Hasegawa-Johnson, “AutoVC: Zero-shot voice style transfer with only autoencoder loss,” in International Conference on Machine Learning , vol. 36, 2019, pp. 5210–5219
2019
Cited alongside, same era.
S. Wisdom, E. Tzinis, H. Erdogan, R. J. Weiss, K. Wilson, and J. R. Hershey, “Unsupervised sound separation using mixture invariant training,” in Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 3846–3857
2020
Cited alongside, same era.
2020
Cited alongside, same era.
K. Qian, Y. Zhang, S. Chang, D. Cox, and M. Hasegawa-Johnson, “Unsupervised speech decomposition via triple information bottleneck,” in International Conference on Machine Learning , vol. 37, 2020, pp. 7836–7846
2020
Cited alongside, same era.
T. Park, J.-Y. Zhu, O. Wang, J. Lu, E. Shechtman, A. A. Efros, and R. Zhang, “Swapping autoencoder for deep image manipulation,” in Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 7198–7211
2020
Cited alongside, same era.
2020
Later among the works it cites.
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “SoundStream: An end-to-end neural audio codec,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 495–507, 2021
2021
Later among the works it cites.
O. Hajihassnai, O. Ardakanian, and H. Khazaei, “ObscureNet: Learning attribute-invariant latent representation for anonymizing sensor data,” in International Conference on Internet-of-Things Design and Implementation , New York, NY, USA, 2021, pp. 40–52
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
H. Yang, K. Zhen, S. Beack, and M. Kim, “Source-aware neural speech coding for noisy speech compression,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 706–710
2021
Later among the works it cites.