Fetching the paper…
Reading the bibliography…
We propose Universal target audio Separation (UniSep), addressing the separation task on arbitrary mixtures of different types of audio.
“Visqol: an objective speech quality model,”
Andrew Hines, Jan Skoglund, Anil C Kokaram, and Naomi Harte, · 2015
Earlier work this paper cites.
“Multichannel audio source separation with deep neural networks,”
Aditya Arie Nugraha, Antoine Liutkus, and Emmanuel Vincent, · 2016
Earlier work this paper cites.
“Audio set: An ontology and human-labeled dataset for audio events,”
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter, · 2017
Earlier work this paper cites.
“Musdb18-a corpus for music separation,”
Zafar Rafii, Antoine Liutkus, Fabian-Robert Stöter, Stylianos Ioannis Mimilakis, and Rachel Bittner, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Supervised speech separation based on deep learning: An overview,”
DeLiang Wang and Jitong Chen, · 2018
Earlier work this paper cites.
“Improving language understanding by generative pre-training,”
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al., · 2018
Earlier work this paper cites.
“Voicefilter: Targeted voice separation by speaker-conditioned spectrogram masking,”
Quan Wang, Hannah Muckenhirn, Kevin Wilson, Prashant Sridhar, Zelin Wu, John Hershey, Rif A Saurous, Ron J Weiss, Ye Jia, and Ignacio Lopez Moreno, · 2018
Earlier work this paper cites.
“Class-conditional embeddings for music source separation,”
Prem Seetharaman, Gordon Wichern, Shrikant Venkataramani, and Jonathan Le Roux, · 2019
Earlier work this paper cites.
“Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,”
Yi Luo and Nima Mesgarani, · 2019
Earlier work this paper cites.
“Universal sound separation,”
Ilya Kavalerov, Scott Wisdom, Hakan Erdogan, Brian Patton, Kevin Wilson, Jonathan Le Roux, and John R. Hershey, · 2019
Earlier work this paper cites.
“Music source separation in the waveform domain,”
Alexandre Défossez, Nicolas Usunier, Léon Bottou, and Francis Bach, · 2019
Earlier work this paper cites.
“Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,”
Kateřina Žmolíková, Marc Delcroix, Keisuke Kinoshita, Tsubasa Ochiai, Tomohiro Nakatani, Lukáš Burget, and Jan Černockỳ, · 2019
Cited alongside, same era.
“Improving universal sound separation using sound classification,”
Efthymios Tzinis, Scott Wisdom, John R Hershey, Aren Jansen, and Daniel PW Ellis, · 2020
Cited alongside, same era.
“Unsupervised sound separation using mixture invariant training,”
Scott Wisdom, Efthymios Tzinis, Hakan Erdogan, Ron Weiss, Kevin Wilson, and John Hershey, · 2020
Cited alongside, same era.
“Libri-light: A benchmark for asr with limited or no supervision,”
Jacob Kahn, Morgane Riviere, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, Julien Karadayi, Vitaliy Liptchinsky, Ronan Collobert, Christian Fuegen, et al., · 2020
Cited alongside, same era.
“Mls: A large-scale multilingual dataset for speech research,”
Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve, and Ronan Collobert, · 2020
“Separate what you describe: Language-queried audio source separation,”
Xubo Liu, Haohe Liu, Qiuqiang Kong, Xinhao Mei, Jinzheng Zhao, Qiushi Huang, Mark D Plumbley, and Wenwu Wang, · 2022
Later among the works it cites.
“Hybrid transformers for music source separation,”
Simon Rouard, Francisco Massa, and Alexandre Défossez, · 2023
Later among the works it cites.
“Gass: Generalizing audio source separation with large-scale data,”
Jordi Pons, Xiaoyu Liu, Santiago Pascual, and Joan Serrà, · 2023
Later among the works it cites.
“Audiolm: a language modeling approach to audio generation,”
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, et al., · 2023
Later among the works it cites.
“Uniaudio: An audio foundation model toward universal audio generation,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Librimix: An open-source dataset for generalizable speech separation,”
Joris Cosentino, Manuel Pariente, Samuele Cornell, Antoine Deleforge, and Emmanuel Vincent, · 2020
Cited alongside, same era.
“Visqol v3: An open source production ready objective speech and audio metric,”
Michael Chinen, Felicia SC Lim, Jan Skoglund, Nikita Gureev, Feargus O’Gorman, and Andrew Hines, · 2020
Cited alongside, same era.
“Exploring the limits of transfer learning with a unified text-to-text transformer,”
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu, · 2020
Cited alongside, same era.
“What’s all the fuss about free universal sound separation data?,”
Scott Wisdom, Hakan Erdogan, Daniel PW Ellis, Romain Serizel, Nicolas Turpault, Eduardo Fonseca, Justin Salamon, Prem Seetharaman, and John R Hershey, · 2021
Cited alongside, same era.
“Soundstream: An end-to-end neural audio codec,”
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi, · 2021
Cited alongside, same era.
“Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,”
Chandan KA Reddy, Vishak Gopal, and Ross Cutler, · 2021
Cited alongside, same era.
“Zero-shot audio source separation through query-based learning from weakly-labeled data,”
Ke Chen, Xingjian Du, Bilei Zhu, Zejun Ma, Taylor Berg-Kirkpatrick, and Shlomo Dubnov, · 2022
Cited alongside, same era.
Dongchao Yang, Jinchuan Tian, Xu Tan, Rongjie Huang, Songxiang Liu, Xuankai Chang, Jiatong Shi, Sheng Zhao, Jiang Bian, Xixin Wu, et al., · 2023
Later among the works it cites.
“High-fidelity audio compression with improved rvqgan,” 2023
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar, · 2023
Later among the works it cites.
Hakan Erdogan, Scott Wisdom, Xuankai Chang, Zalán Borsos, Marco Tagliasacchi, Neil Zeghidour, and John R Hershey, · 2023
Later among the works it cites.
“Universal source separation with weakly labelled data,”
Qiuqiang Kong, Ke Chen, Haohe Liu, Xingjian Du, Taylor Berg-Kirkpatrick, Shlomo Dubnov, and Mark D Plumbley, · 2023
Later among the works it cites.
“Separate anything you describe,”
Xubo Liu, Qiuqiang Kong, Yan Zhao, Haohe Liu, Yi Yuan, Yuzhuo Liu, Rui Xia, Yuxuan Wang, Mark D Plumbley, and Wenwu Wang, · 2023
Later among the works it cites.
“Consistent and relevant: Rethink the query embedding in general sound separation,”
Yuanyuan Wang, Hangting Chen, Dongchao Yang, Jianwei Yu, Chao Weng, Zhiyong Wu, and Helen Meng, · 2024
Later among the works it cites.
“Clapsep: Leveraging contrastive pre-trained model for multi-modal query-conditioned target sound extraction,”
Hao Ma, Zhiyuan Peng, Xu Li, Mingjie Shao, Xixin Wu, and Ju Liu, · 2024
Later among the works it cites.
“Soloaudio: Target sound extraction with language-oriented audio diffusion transformer,”
Helin Wang, Jiarui Hai, Yen-Ju Lu, Karan Thakkar, Mounya Elhilali, and Najim Dehak, · 2024
Later among the works it cites.