Fetching the paper…
Reading the bibliography…
Deep-learning based speech separation models confront poor generalization problem that even the state-of-the-art models could abruptly fail when evaluating them in mismatch conditions.
“On the uniform convergence of relative frequencies of events to their probabilities,”
V. Vapnik and A. Y. Chervonenkis, · 1971
Earlier work this paper cites.
“Acceleration of stochastic approximation by averaging,”
B. T. Polyak and A. B. Juditsky, · 1992
Earlier work this paper cites.
“Statistical learning theory,” 1998
Vladimir Vapnik, · 1998
Earlier work this paper cites.
“Semi-supervised learning (chapelle, o. et al., eds.; 2006)[book reviews],”
Olivier Chapelle, Bernhard Scholkopf, and Alexander Zien, · 2009
Earlier work this paper cites.
“Vocal tract length perturbation (vtlp) improves speech recognition,”
Navdeep Jaitly and Geoffrey E Hinton, · 2013
Earlier work this paper cites.
“Generative adversarial nets,”
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Semi-supervised learning with ladder networks,”
Antti Rasmus, Mathias Berglund, Mikko Honkala, Harri Valpola, and Tapani Raiko, · 2015
Earlier work this paper cites.
“Audio augmentation for speech recognition,”
Tom Ko, Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Single-channel multi-speaker separation using deep clustering,”
Yusuf Isik, Jonathan Le Roux, Zhuo Chen, Shinji Watanabe, and John R Hershey, · 2016
Earlier work this paper cites.
“Regularization with stochastic transformations and perturbations for deep semi-supervised learning,”
Mehdi Sajjadi, Mehran Javanmardi, and Tolga Tasdizen, · 2016
Earlier work this paper cites.
“Temporal ensembling for semi-supervised learning,”
Samuli Laine and Timo Aila, · 2016
Cited alongside, same era.
“Deep clustering: Discriminative embeddings for segmentation and separation,”
John R Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe, · 2016
Cited alongside, same era.
“Deep attractor network for single-microphone speaker separation,”
Zhuo Chen, Yi Luo, and Nima Mesgarani, · 2017
Cited alongside, same era.
“Permutation invariant training of deep models for speaker-independent multi-talker speech separation,”
Dong Yu, Morten Kolbæk, Zheng-Hua Tan, and Jesper Jensen, · 2017
Cited alongside, same era.
“Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,”
Morten Kolbæk, Dong Yu, Zheng-Hua Tan, Jesper Jensen, Morten Kolbaek, Dong Yu, Zheng-Hua Tan, and Jesper Jensen, · 2017
Cited alongside, same era.
“Single channel speech separation with constrained utterance level permutation invariant training using grid lstm,”
Chenglin Xu, Wei Rao, Xiong Xiao, Eng Siong Chng, and Haizhou Li, · 2018
Later among the works it cites.
“mixup: Beyond empirical risk minimization,”
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz, · 2018
Later among the works it cites.
“Virtual adversarial training: a regularization method for supervised and semi-supervised learning,”
Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, and Shin Ishii, · 2018
Later among the works it cites.
“Smooth neighbors on teacher graphs for semi-supervised learning,”
Yucen Luo, Jun Zhu, Mengxi Li, Yong Ren, and Bo Zhang, · 2018
Later among the works it cites.
“Alternative objective functions for deep clustering,”
Zhong-Qiu Wang, Jonathan Le Roux, and John R Hershey, · 2018
Later among the works it cites.
“End-to-end speech separation with unfolded iterative phase reconstruction,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results,”
Antti Tarvainen and Harri Valpola, · 2017
Cited alongside, same era.
“Understanding deep learning requires rethinking generalization,”
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals, · 2017
Cited alongside, same era.
“Smiles enumeration as data augmentation for neural network modeling of molecules,”
Esben Jannik Bjerrum, · 2017
Cited alongside, same era.
“Data augmentation generative adversarial networks,”
Antreas Antoniou, Amos Storkey, and Harrison Edwards, · 2017
Cited alongside, same era.
“Deep extractor network for target speaker recovery from single channel speech mixtures,”
Jun Wang, Jie Chen, Dan Su, Lianwu Chen, Meng Yu, Yanmin Qian, and Dong Yu, · 2018
Cited alongside, same era.
“Speaker-independent speech separation with deep attractor network,”
Yi Luo, Zhuo Chen, and Nima Mesgarani, · 2018
Cited alongside, same era.
Zhong-Qiu Wang, Jonathan Le Roux, DeLiang Wang, and John R Hershey, · 2018
Later among the works it cites.
“Real-time single-channel dereverberation and separation with time-domain audio separation network.,”
Yi Luo and Nima Mesgarani, · 2018
Later among the works it cites.
“Conv-tasnet: Surpassing ideal time-frequency masking for speech separation,”
Yi Luo and Nima Mesgarani, · 2019
Closest in time.
“Interpolation consistency training for semi-supervised learning,”
Vikas Verma, Alex Lamb, Juho Kannala, Yoshua Bengio, and David Lopez-Paz, · 2019
Closest in time.
“Improved speech separation with time-and-frequency cross-domain joint embedding and clustering,”
Gene-Ping Yang, Chao-I Tuan, Hung-Yi Lee, and Lin-shan Lee, · 2019
Closest in time.