Fetching the paper…
Reading the bibliography…
Language-queried audio source separation (LASS) is a new paradigm for computational auditory scene analysis (CASA).
S. R. Quackenbush, T. P. Barnwell, and M. A. Clements, “Objective measures of speech quality,” Prentice Hall , 1988
1988
Earlier work this paper cites.
A. J. Bell and T. J. Sejnowski, “An information-maximization approach to blind separation and blind deconvolution,” Neural computation , vol. 7, no. 6, pp. 1129–1159, 1995
1995
Earlier work this paper cites.
P. Scalart et al. , “Speech enhancement based on a priori signal to noise estimation,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , vol. 2, 1996, pp. 629–632
1996
Earlier work this paper cites.
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Perceptual evaluation of speech quality (PESQ) - a new method for speech quality assessment of telephone networks and codecs,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP). , vol. 2, 2001, pp. 749–752
2001
Earlier work this paper cites.
S. Haykin and Z. Chen, “The cocktail party problem,” Neural Computation , vol. 17, no. 9, pp. 1875–1902, 2005
2005
Earlier work this paper cites.
D. Wang and G. J. Brown, Computational auditory scene analysis: Principles, algorithms, and applications . Wiley-IEEE press, 2006
2006
Earlier work this paper cites.
Y. Hu and P. C. Loizou, “Evaluation of objective quality measures for speech enhancement,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 16, no. 1, pp. 229–238, 2007
2007
Earlier work this paper cites.
S. Rubin, F. Berthouzoz, G. J. Mysore, W. Li, and M. Agrawala, “Content-based tools for editing audio stories,” in Proceedings of the 26th Annual ACM Symposium on User Interface Software and Technology , 2013, pp. 113–122
2013
Earlier work this paper cites.
A. Narayanan and D. Wang, “Ideal ratio mask estimation using deep neural networks for robust speech recognition,” in International Conference on Acoustics, Speech and Signal Processing . IEEE, 2013, pp. 7092–7096
2013
Earlier work this paper cites.
C. Veaux, J. Yamagishi, and S. King, “The voice bank corpus: Design, collection and data analysis of a large regional accent speech database,” in International Conference Oriental COCOSDA with Conference on Asian Spoken Language Research and Evaluation (O-COCOSDA/CASLRE) , 2013, pp. 1–4
2013
Earlier work this paper cites.
J. Thiemann, N. Ito, and E. Vincent, “DEMAND: a collection of multi-channel recordings of acoustic noise in diverse environments,” in Proceedings of Meetings on Acoustics , 2013, pp. 1–6
2013
Earlier work this paper cites.
K. J. Piczak, “ESC: Dataset for environmental sound classification,” in Proceedings of the 23rd ACM International Conference on Multimedia , 2015, pp. 1015–1018
2015
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional networks for biomedical image segmentation,” in Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18 , 2015, pp. 234–241
2015
Earlier work this paper cites.
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” in International Conference on Acoustics, Speech and Signal Processing . IEEE, 2016, pp. 31–35
2016
Earlier work this paper cites.
S. Pascual, A. Bonafonte, and J. Serra, “SEGAN: Speech enhancement generative adversarial network,” in INTERSPEECH , 2017
2017
Earlier work this paper cites.
Y. Peng, X. Huang, and Y. Zhao, “An overview of cross-media retrieval: Concepts, methodologies, benchmarks, and challenges,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 28, no. 9, pp. 2372–2385, 2017
2017
Earlier work this paper cites.
D. Yu, M. Kolbæk, Z.-H. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 241–245
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 776–780
2017
Earlier work this paper cites.
K. Drossos, S. Adavanne, and T. Virtanen, “Automated audio captioning with recurrent neural networks,” in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) , 2017, pp. 374–378
2017
Earlier work this paper cites.
D. Wang and J. Chen, “Supervised speech separation based on deep learning: An overview,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 10, pp. 1702–1726, 2018
2018
Earlier work this paper cites.
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba, “The sound of pixels,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 570–586
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Z.-Q. Wang, J. Le Roux, and J. R. Hershey, “Multi-channel deep clustering: Discriminative spectral and spatial embeddings for speaker-independent speech separation,” in International Conference on Acoustics, Speech and Signal Processing . IEEE, 2018, pp. 1–5
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville, “FiLm: Visual reasoning with a general conditioning layer,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
I. Kavalerov, S. Wisdom, H. Erdogan, B. Patton, K. Wilson, J. Le Roux, and J. R. Hershey, “Universal sound separation,” in 2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) , 2019, pp. 175–179
2019
Earlier work this paper cites.
Y. Luo and N. Mesgarani, “Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation,” IEEE/ACM transactions on Audio, Speech, and Language Processing , vol. 27, no. 8, pp. 1256–1266, 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
J. Wu, Y. Xu, S.-X. Zhang, L.-W. Chen, M. Yu, L. Xie, and D. Yu, “Time domain audio visual speech separation,” in IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , 2019, pp. 667–673
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Y. Wang, D. Stoller, R. M. Bittner, and J. P. Bello, “Few-shot musical source separation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022, pp. 121–125
2022
Later among the works it cites.
M. Delcroix, J. B. Vázquez, T. Ochiai, K. Kinoshita, Y. Ohishi, and S. Araki, “Soundbeam: Target sound extraction conditioned on sound-class labels and enrollment clues for increased performance and continuous learning,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 31, pp. 121–136, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
K. Kilgour, B. Gfeller, Q. Huang, A. Jansen, S. Wisdom, and M. Tagliasacchi, “Text-driven separation of arbitrary sounds,” in INTERSPEEH , 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. D. Kim, B. Kim, H. Lee, and G. Kim, “Audiocaps: Generating captions for audios in the wild,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , 2019, pp. 119–132
2019
Cited alongside, same era.
J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “SDR–half-baked or well done?” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 626–630
2019
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Z.-Q. Wang, P. Wang, and D. Wang, “Complex spectral mapping for single-and multi-channel speech enhancement and robust asr,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 1778–1787, 2020
2020
Cited alongside, same era.
S. Wisdom, E. Tzinis, H. Erdogan, R. Weiss, K. Wilson, and J. Hershey, “Unsupervised sound separation using mixture invariant training,” Advances in Neural Information Processing Systems (NeurIPS) , vol. 33, pp. 3846–3857, 2020
2020
Cited alongside, same era.
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, “PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 2880–2894, 2020
2020
Cited alongside, same era.
H. Chen, W. Xie, A. Vedaldi, and A. Zisserman, “VGGSound: A large-scale audio-visual dataset,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 721–725
2020
Cited alongside, same era.
Y. Xiao, X. Liu, J. King, A. Singh, E. S. Chng, M. D. Plumbley, and W. Wang, “Continual learning for on-device environmental sound classification,” in Proceedings of the 7th Detection and Classification of Acoustic Scenes and Events (DCASE) Workshop , 2022
2022
Later among the works it cites.
X. Mei, X. Liu, J. Sun, M. Plumbley, and W. Wang, “On metric learning for audio-text cross-modal retrieval,” in Proc. Interspeech , 2022, pp. 4142–4146
2022
Later among the works it cites.
2022
Later among the works it cites.
K. Chen, X. Du, B. Zhu, Z. Ma, T. Berg-Kirkpatrick, and S. Dubnov, “HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022, pp. 646–650
2022
Later among the works it cites.
2022
Later among the works it cites.
V. W. Liang, Y. Zhang, Y. Kwon, S. Yeung, and J. Y. Zou, “Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning,” Advances in Neural Information Processing Systems , vol. 35, pp. 17 612–17 625, 2022
2022
Later among the works it cites.
2023
Closest in time.
B. Veluri, J. Chan, M. Itani, T. Chen, T. Yoshioka, and S. Gollakota, “Real-time target sound extraction,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2023, pp. 1–5
2023
Closest in time.
H.-W. Dong, N. Takahashi, Y. Mitsufuji, J. McAuley, and T. Berg-Kirkpatrick, “CLIPSep: Learning text-queried sound separation with noisy unlabeled videos,” in Proceedings of International Conference on Learning Representations (ICLR) , 2023
2023
Closest in time.
2023
Closest in time.
L. Wan, H. Liu, Y. Zhou, and J. Jia, “Multi-loss convolutional network with time-frequency attention for speech enhancement,” in International Conference on Information Communication and Signal Processing . IEEE, 2023, pp. 629–633
2023
Closest in time.
E. Tzinis, G. Wichern, P. Smaragdis, and J. Le Roux, “Optimal condition training for target source separation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2023, pp. 1–5
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
D. Yang, J. Yu, H. Wang, W. Wang, C. Weng, Y. Zou, and D. Yu, “Diffsound: Discrete diffusion model for text-to-sound generation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2023
2023
Closest in time.
Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2023, pp. 1–5
2023
Closest in time.
J. Liang, X. Liu, H. Liu, H. Phan, E. Benetos, M. D. Plumbley, and W. Wang, “Adapting language-audio models as few-shot audio learners,” in INTERSPEEH , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2024
Closest in time.
S.-W. Fu, C.-F. Liao, Y. Tsao, and S.-D. Lin, “Metricgan: Generative adversarial networks based black-box metric scores optimization for speech enhancement,” in International Conference on Machine Learning . PMLR, 2019, pp. 2031–2041
2041
Closest in time.