Fetching the paper…
Reading the bibliography…
The problem of speech separation, also known as the cocktail party problem, refers to the task of isolating a single speech signal from a mixture of speech signals.
CSR-I (WSJ0) complete, 1993
Garofolo, John S., Graff, David, Paul, Doug, and Pallett, David · 1993
Earlier work this paper cites.
Elements of Information Theory (Wiley Series in Telecommunications and Signal Processing)
Thomas M. Cover and Joy A. Thomas · 2006
Earlier work this paper cites.
What makes a good model of natural images?
Yair Weiss and William T Freeman · 2007
Earlier work this paper cites.
MNIST handwritten digit database
Yann LeCun and Corinna Cortes · 2010
Earlier work this paper cites.
A survey on single channel speech separation
G Logeshwari and GS Anandha Mala · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop
Fisher Yu, Ari Seff, Yinda Zhang, Shuran Song, Thomas Funkhouser, and Jianxiong Xiao · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Deep clustering: Discriminative embeddings for segmentation and separation
John R Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe · 2016
Earlier work this paper cites.
A kernelized stein discrepancy for goodness-of-fit tests
Qiang Liu, Jason Lee, and Michael Jordan · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks
Morten Kolbæk, Dong Yu, Zheng-Hua Tan, and Jesper Jensen · 2017
Cited alongside, same era.
Permutation invariant training of deep models for speaker-independent multi-talker speech separation
Dong Yu, Morten Kolbæk, Zheng-Hua Tan, and Jesper Jensen · 2017
Cited alongside, same era.
Efficient neural audio synthesis
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron Oord, Sander Dieleman, and Koray Kavukcuoglu · 2018
Cited alongside, same era.
Tasnet: time-domain audio separation network for real-time, single-channel speech separation
Yi Luo and Nima Mesgarani · 2018
Wavegrad: Estimating gradients for waveform generation
Nanxin Chen, Yu Zhang, Heiga Zen, Ron J Weiss, Mohammad Norouzi, and William Chan · 2020
Later among the works it cites.
Source separation with deep generative priors
Vivek Jayaram and John Thickstun · 2020
Later among the works it cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae · 2020
Later among the works it cites.
Diffwave: A versatile diffusion model for audio synthesis
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro · 2020
Later among the works it cites.
Voice separation with an unknown number of multiple speakers
Eliya Nachmani, Yossi Adi, and Lior Wolf · 2020
Later among the works it cites.
Furcanext: End-to-end monaural speech separation with dynamic gated dilated temporal convolutional networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Single-channel speech presence probability estimation and noise tracking
Rainer Martin and Israel Cohen · 2018
Cited alongside, same era.
High fidelity speech synthesis with adversarial networks
Mikołaj Bińkowski, Jeff Donahue, Sander Dieleman, Aidan Clark, Erich Elsen, Norman Casagrande, Luis C Cobo, and Karen Simonyan · 2019
Cited alongside, same era.
SDR–half-baked or well done?
Jonathan Le Roux, Scott Wisdom, Hakan Erdogan, and John R Hershey · 2019
Cited alongside, same era.
Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech separation
Yi Luo, Zhuo Chen, and Takuya Yoshioka · 2019
Cited alongside, same era.
Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation
Yi Luo and Nima Mesgarani · 2019
Cited alongside, same era.
Generative modeling by estimating gradients of the data distribution
Yang Song and Stefano Ermon · 2019
Cited alongside, same era.
Liwen Zhang, Ziqiang Shi, Jiqing Han, Anyan Shi, and Ding Ma · 2020
Later among the works it cites.
Wavegrad 2: Iterative refinement for text-to-speech synthesis
Nanxin Chen, Yu Zhang, Heiga Zen, Ron J Weiss, Mohammad Norouzi, Najim Dehak, and William Chan · 2021
Later among the works it cites.
Many-speakers single channel speech separation with optimal permutation training
Shaked Dovrat, Eliya Nachmani, and Lior Wolf · 2021
Later among the works it cites.
Attention is all you need in speech separation
Cem Subakan, Mirco Ravanelli, Samuele Cornell, Mirko Bronzi, and Jianyuan Zhong · 2021
Later among the works it cites.
Sepit: Approaching a single channel speech separation bound
Shahar Lutati, Eliya Nachmani, and Lior Wolf · 2022
Later among the works it cites.
Diffusion-based generative speech source separation
Robin Scheibler, Youna Ji, Soo-Whan Chung, Jaeuk Byun, Soyeon Choe, and Min-Seok Choi · 2022
Later among the works it cites.