Fetching the paper…
Reading the bibliography…
Universal source separation targets at separating the audio sources of an arbitrary mix, removing the constraint to operate on a specific domain like speech or music.
“TIMIT Acoustic Phonetic Continuous Speech Corpus,”
J. S. Garofolo et al., · 1993
Earlier work this paper cites.
“ENST-Drums: An Extensive Audio-Visual Database for Drum Signals Processing,”
O. Gillet and G. Richard, · 2006
Earlier work this paper cites.
“The Diverse Environments Multi-Channel Acoustic Noise Database (DEMAND): A Database of Multichannel Environmental Noise Recordings,”
J. Thiemann, N. Ito, and E. Vincent, · 2013
Earlier work this paper cites.
“Can We Automatically Transform Speech Recorded on Common Consumer Devices in Real-World Environments into Professional Production Quality Speech?—A Dataset, Insights, and Challenges,”
G. J. Mysore, · 2014
Earlier work this paper cites.
“Automatic Identification of Emotional Cues in Chinese Opera Singing,”
D. A. Black, M. Li, and M. Tian, · 2014
Earlier work this paper cites.
“Librispeech: An ASR Corpus Based on Public Domain Audio Books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit,”
C. Veaux, J. Yamagishi, and K. MacDonald, · 2017
Earlier work this paper cites.
“Multitalker Speech Separation with Utterance-level Permutation Invariant Training of Deep Recurrent Neural Networks,”
M. Kolbæk, D. Yu, Z. Tan, and J. Jensen, · 2017
Earlier work this paper cites.
“Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation,”
A. Ephrat, I. Mosseri, O. Lang, T. Dekel, K. Wilson, A. Hassidim, W. T. Freeman, and M. Rubinstein, · 2018
Earlier work this paper cites.
“VocalSet: A Singing Voice Dataset,”
J. Wilkins, P. Seetharaman, A. Wahl, and B. Pardo, · 2018
Earlier work this paper cites.
“The Sound of Pixels,”
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba, · 2018
Earlier work this paper cites.
“The 2018 Signal Separation Evaluation Campaign,”
F. Stöter, A. Liutkus, and N. Ito, · 2018
Earlier work this paper cites.
“Universal Sound Separation,”
I. Kavalerov, S. Wisdom, H. Erdogan, B. Patton, K. Wilson, J. Le Roux, and J. R. Hershey, · 2019
Cited alongside, same era.
“WHAM!: Extending Speech Separation to Noisy Environments,”
G. Wichern, J. Antognini, M. Flynn, L. R. Zhu, E. McQuinn, D. Crow, E. Manilow, and J. Le Roux, · 2019
Cited alongside, same era.
“Cutting Music Source Separation Some Slakh: A Dataset to Study the Impact of Training Data Quality and Quantity,”
E. Manilow, G. Wichern, P. Seetharaman, and J. Le Roux, · 2019
Cited alongside, same era.
“MUSDB18-HQ - An Uncompressed Version of MUSDB18,” 2019
Z. Rafii, A. Liutkus, F. Stöter, S. I. Mimilakis, and R. Bittner, · 2019
Cited alongside, same era.
“SDR–Half-Baked or Well Done?,”
J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, · 2019
Cited alongside, same era.
“Improving Universal Sound Separation Using Sound Classification,”
E. Tzinis, S. Wisdom, J. R. Hershey, A. Jansen, and D. P. W. Ellis, · 2020
“FSD50K: An Open Dataset of Human-Labeled Sound Events,”
E. Fonseca, X. Favory, J. Pons, F. Font, and X. Serra, · 2021
Later among the works it cites.
“Compute and Memory Efficient Universal Sound Source Separation,”
E. Tzinis, Z. Wang, X. Jiang, and P. Smaragdis, · 2022
Later among the works it cites.
“EGFxSet: Electric Guitar Tones Processed Through Real Effects of Distortion, Modulation, Delay and Reverb,”
H. Pedroza, G. Meza, and I. R. Roman, · 2022
Later among the works it cites.
“The Cocktail Fork Problem: Three-Stem Audio Separation for Real-world Soundtracks,”
D. Petermann, G. Wichern, Z. Wang, and J. Le Roux, · 2022
Later among the works it cites.
“An Efficient Encoder-Decoder Architecture with Top-down Attention for Speech Separation,”
K. Li, R. Yang, and X. Hu, · 2023
Closest in time.
“Hybrid Transformers for Music Source Separation,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Unsupervised Sound Separation Using Mixture Invariant Training,”
S. Wisdom, E. Tzinis, H. Erdogan, R. Weiss, K. Wilson, and J. Hershey, · 2020
Cited alongside, same era.
“An Empirical Study of Conv-TasNet,”
B. Kadıoğlu, M. Horgan, X. Liu, J. Pons, D. Darcy, and V. Kumar, · 2020
Cited alongside, same era.
“LibriMix: An Open-Source Dataset for Generalizable Speech Separation,”
J. Cosentino, M. Pariente, S. Cornell, A. Deleforge, and E. Vincent, · 2020
Cited alongside, same era.
“What’s All the Fuss About Free Universal Sound Separation Data?,”
S. Wisdom, H. Erdogan, D. P. W. Ellis, R. Serizel, N. Turpault, E. Fonseca, J. Salamon, P. Seetharaman, and J. R. Hershey, · 2021
Cited alongside, same era.
“Sparse, Efficient, and Semantic Mixture Invariant Training: Taming In-the-wild Unsupervised Sound Separation,”
S. Wisdom, A. Jansen, R. J. Weiss, H. Erdogan, and J. R. Hershey, · 2021
Cited alongside, same era.
“HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,”
W. Hsu, B. Bolte, Y. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, · 2021
Cited alongside, same era.
S. Rouard, F. Massa, and A. Défossez, · 2023
Closest in time.
“Music Source Separation With Band-Split RNN,”
Y. Luo and J. Yu, · 2023
Closest in time.
“Adversarial Permutation Invariant Training for Universal Sound Separation,”
E. Postolache, J. Pons, S. Pascual, and J. Serrà, · 2023
Closest in time.
“Universal Source Separation with Weakly Labelled Data,”
Q. Kong, K. Chen, H. Liu, X. Du, T. Berg-Kirkpatrick, S. Dubnov, and M. D. Plumbley, · 2023
Closest in time.
“Separate Anything You Describe,”
X. Liu, Q. Kong, Y. Zhao, H. Liu, Y. Yuan, Y. Liu, R. Xia, Y. Wang, M. D. Plumbley, and W. Wang, · 2023
Closest in time.
“CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled Videos,”
H. Dong, N. Takahashi, Y. Mitsufuji, J. McAuley, and T. Berg-Kirkpatrick, · 2023
Closest in time.
“High Fidelity Speech Enhancement with Band-split RNN,”
J. Yu, H. Chen, Y. Luo, R. Gu, and C. Weng, · 2023
Closest in time.