Fetching the paper…
Reading the bibliography…
In daily life, we encounter a variety of sounds, both desirable and undesirable, with limited control over their presence and volume.
E. C. Cherry, “Some experiments on the recognition of speech, with one and with two ears,” The Journal of the acoustical society of America , vol. 25, no. 5, pp. 975–979, 1953
1953
Earlier work this paper cites.
S. Haykin and Z. Chen, “The cocktail party problem,” Neural computation , vol. 17, no. 9, pp. 1875–1902, 2005
2005
Earlier work this paper cites.
J. Kates, Digital Hearing Aids . Plural Publishing, Incorporated, 2008. [Online]. Available: https://books.google.com/books?id=xDI7CQAAQBAJ
2008
Earlier work this paper cites.
J. H. McDermott, “The cocktail party problem,” Current Biology , vol. 19, no. 22, pp. R1024–R1027, 2009
2009
Earlier work this paper cites.
N. J. Bryan, G. J. Mysore, and G. Wang, “Source separation of polyphonic music with interactive user-feedback on a piano roll display.” in ISMIR , 2013, pp. 119–124
2013
Earlier work this paper cites.
J. L. Clark and D. W. Swanepoel, “Technology for hearing loss–as we know it, and as we dream it,” Disability and Rehabilitation: Assistive Technology , vol. 9, no. 5, pp. 408–413, 2014
2014
Earlier work this paper cites.
J. R. Hershey, Z. Chen, J. L. Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
S. Launer, J. A. Zakis, and B. C. J. Moore, “Hearing aid signal processing,” in Hearing Aids , ser. Springer Handbook of Auditory Research. Springer, 2016, vol. 56, pp. 93–130. [Online]. Available: https://link.springer.com/chapter/10.1007/978-3-319-33036-5_4
2016
Earlier work this paper cites.
S. Van Eyndhoven, T. Francart, and A. Bertrand, “Eeg-informed attended speaker extraction from recorded speech mixtures with application in neuro-steered hearing prostheses,” IEEE Transactions on Biomedical Engineering , vol. 64, no. 5, pp. 1045–1056, 2016
2016
Earlier work this paper cites.
C. Lea, R. Vidal, A. Reiter, and G. D. Hager, “Temporal convolutional networks: A unified approach to action segmentation,” in Computer Vision–ECCV 2016 Workshops: Amsterdam, The Netherlands, October 8-10 and 15-16, 2016, Proceedings, Part III 14 . Springer, 2016, pp. 47–54
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, Eds., vol. 30. Curran Associates, Inc., 2017. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf
2017
Earlier work this paper cites.
F. Chollet, “Xception: Deep learning with depthwise separable convolutions,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 1251–1258
2017
Earlier work this paper cites.
D. Yu, M. Kolbæk, Z.-H. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 241–245
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 776–780
2017
Earlier work this paper cites.
S. Settle, J. Le Roux, T. Hori, S. Watanabe, and J. R. Hershey, “End-to-end multi-speaker speech recognition,” in 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2018, pp. 4819–4823
2018
Earlier work this paper cites.
T. Afouras, J. S. Chung, and A. Zisserman, “The conversation: Deep audio-visual speech enhancement,” in Proc. Interspeech 2018 , 2018, pp. 3244–3248
2018
Earlier work this paper cites.
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba, “The sound of pixels,” in The European Conference on Computer Vision (ECCV) , September 2018
2018
Earlier work this paper cites.
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville, “FiLM: Visual reasoning with a general conditioning layer,” in Proceedings of the AAAI conference on artificial intelligence , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
J. L. Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “SDR – half-baked or well done?” ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 626–630, 2018. [Online]. Available: https://api.semanticscholar.org/CorpusID:53246666
2018
Earlier work this paper cites.
C. Han, J. O’Sullivan, Y. Luo, J. Herrero, A. D. Mehta, and N. Mesgarani, “Speaker-independent auditory attention decoding without access to clean speech sources,” Science advances , vol. 5, no. 5, p. eaav6134, 2019
2019
Earlier work this paper cites.
Y. Luo and N. Mesgarani, “Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 8, p. 1256–1266, Aug. 2019. [Online]. Available: http://dx.doi.org/10.1109/TASLP.2019.2915167
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
I. Kavalerov, S. Wisdom, H. Erdogan, B. Patton, K. Wilson, J. Le Roux, and J. R. Hershey, “Universal sound separation,” in 2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) , 2019, pp. 175–179
2019
Cited alongside, same era.
K. Žmolíková, M. Delcroix, K. Kinoshita, T. Ochiai, T. Nakatani, L. Burget, and J. Černocký, “Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,” IEEE Journal of Selected Topics in Signal Processing , vol. 13, no. 4, pp. 800–814, 2019
2019
Cited alongside, same era.
Q. Wang, H. Muckenhirn, K. Wilson, P. Sridhar, Z. Wu, J. R. Hershey, R. A. Saurous, R. J. Weiss, Y. Jia, and I. L. Moreno, “VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking,” in Proc. Interspeech 2019 , 2019, pp. 2728–2732. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2019-1101
2019
Cited alongside, same era.
Y. Luo and J. Yu, “Music source separation with band-split rnn,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 31, pp. 1893–1901, 2023
2023
Later among the works it cites.
K. Zmolikova, M. Delcroix, T. Ochiai, K. Kinoshita, J. Černocký, and D. Yu, “Neural target speech extraction: An overview,” IEEE Signal Processing Magazine , vol. 40, no. 3, p. 8–29, May 2023. [Online]. Available: http://dx.doi.org/10.1109/MSP.2023.3240008
2023
Later among the works it cites.
2023
Later among the works it cites.
F. Kreuk, G. Synnaeve, A. Polyak, U. Singer, A. Défossez, J. Copet, D. Parikh, Y. Taigman, and Y. Adi, “Audiogen: Textually guided audio generation,” in The Eleventh International Conference on Learning Representations , 2023. [Online]. Available: https://openreview.net/forum?id=CYK7RfcOzQ4
2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Gu, L. Chen, S.-X. Zhang, J. Zheng, Y. Xu, M. Yu, D. Su, Y. Zou, and D. Yu, “Neural spatial filter: Target speaker speech separation assisted with directional information,” in Interspeech , 2019. [Online]. Available: https://api.semanticscholar.org/CorpusID:197629146
2019
Cited alongside, same era.
J. Heitkaemper, T. Fehér, M. J. Freitag, and R. Häb-Umbach, “A study on online source extraction in the presence of changing speaker positions,” in International Conference on Statistical Language and Speech Processing , 2019. [Online]. Available: https://api.semanticscholar.org/CorpusID:203565755
2019
Cited alongside, same era.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Cited alongside, same era.
N. Turpault, S. Wisdom, H. Erdogan, J. R. Hershey, R. Serizel, E. Fonseca, P. Seetharaman, and J. Salamon, “Improving Sound Event Detection In Domestic Environments Using Sound Separation,” in DCASE Workshop 2020 - Detection and Classification of Acoustic Scenes and Events , Tokyo / Virtual, Japan, Nov. 2020. [Online]. Available: https://inria.hal.science/hal-02891700
2020
Cited alongside, same era.
H. Chen, W. Xie, A. Vedaldi, and A. Zisserman, “VGGSound: A large-scale audio-visual dataset,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 721–725
2020
Cited alongside, same era.
C. Subakan, M. Ravanelli, S. Cornell, M. Bronzi, and J. Zhong, “Attention is all you need in speech separation,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 21–25
2021
Cited alongside, same era.
S. Wisdom, H. Erdogan, D. P. W. Ellis, R. Serizel, N. Turpault, E. Fonseca, J. Salamon, P. Seetharaman, and J. R. Hershey, “What’s all the fuss about free universal sound separation data?” in ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 186–190
2021
Cited alongside, same era.
P. Ma, S. Petridis, and M. Pantic, “End-to-end audio-visual speech recognition with conformers,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 7613–7617
2021
Cited alongside, same era.
E. Tzinis, S. Wisdom, A. Jansen, S. Hershey, T. Remez, D. Ellis, and J. R. Hershey, “Into the wild with audioscope: Unsupervised audio-visual separation of on-screen sounds,” in International Conference on Learning Representations , 2021. [Online]. Available: https://openreview.net/forum?id=MDsQkFP1Aw
2021
Cited alongside, same era.
Later among the works it cites.
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley, “AudioLDM: Text-to-audio generation with latent diffusion models,” Proceedings of the International Conference on Machine Learning , 2023
2023
Later among the works it cites.
Y. Wang, Z. Ju, X. Tan, L. He, Z. Wu, J. Bian, and S. Zhao, “Audit: Audio editing by following instructions with latent diffusion models,” in NeurIPS 2023 , December 2023. [Online]. Available: https://www.microsoft.com/en-us/research/publication/audit-audio-editing-by-following-instructions-with-latent-diffusion-models/
2023
Later among the works it cites.
M. Le, A. Vyas, B. Shi, B. Karrer, L. Sari, R. Moritz, M. Williamson, V. Manohar, Y. Adi, J. Mahadeokar, and W.-N. Hsu, “Voicebox: Text-guided multilingual universal speech generation at scale,” in Thirty-seventh Conference on Neural Information Processing Systems , 2023. [Online]. Available: https://openreview.net/forum?id=gzCS252hCO
2023
Later among the works it cites.
2023
Later among the works it cites.
B. Elizalde, S. Deshmukh, M. Al Ismail, and H. Wang, “Clap learning audio concepts from natural language supervision,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
A. Gu and T. Dao, “Mamba: Linear-time sequence modeling with selective state spaces,” in First Conference on Language Modeling , 2024. [Online]. Available: https://openreview.net/forum?id=tEYskw1VY2
2024
Closest in time.
M. Elminshawi, W. Mack, S. R. Chetupalli, S. Chakrabarty, and E. A. P. Habets, “New insights on the role of auxiliary information in target speaker extraction,” Frontiers in Signal Processing , vol. Volume 4 - 2024, 2024. [Online]. Available: https://www.frontiersin.org/journals/signal-processing/articles/10.3389/frsip.2024.1440401
2024
Closest in time.
S. Ji, J. Zuo, M. Fang, Z. Jiang, F. Chen, X. Duan, B. Huai, and Z. Zhao, “Textrolspeech: A text style control speech corpus with codec language text-to-speech models,” in ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2024, pp. 10 301–10 305
2024
Closest in time.
M. Yang, C. Zhang, Y. Xu, Z. Xu, H. Wang, B. Raj, and D. Yu, “usee: Unified speech enhancement and editing with conditional diffusion models,” in ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2024, pp. 7125–7129
2024
Closest in time.
B. Han, J. Dai, W. Hao, X. He, D. Guo, J. Chen, Y. Wang, Y. Qian, and X. Song, “Instructme: An instruction guided music edit framework with latent diffusion models,” in Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, IJCAI-24 , K. Larson, Ed. International Joint Conferences on Artificial Intelligence Organization, 8 2024, pp. 5835–5843, main Track. [Online]. Available: https://doi.org/10.24963/ijcai.2024/645
2024
Closest in time.
R. Huang, M. Li, D. Yang, J. Shi, X. Chang, Z. Ye, Y. Wu, Z. Hong, J. Huang, J. Liu et al. , “Audiogpt: Understanding and generating speech, music, sound, and talking head,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 21, 2024, pp. 23 802–23 804
2024
Closest in time.
J. Liang, H. Zhang, H. Liu, Y. Cao, Q. Kong, X. Liu, W. Wang, M. D. Plumbley, H. Phan, and E. Benetos, “Wavcraft: Audio editing and generation with large language models,” in ICLR 2024 Workshop on Large Language Model (LLM) Agents , 2024. [Online]. Available: https://openreview.net/forum?id=xJw7x2ZBex
2024
Closest in time.
L. Zhu, B. Liao, Q. Zhang, X. Wang, W. Liu, and X. Wang, “Vision mamba: Efficient visual representation learning with bidirectional state space model,” in Proceedings of the 41st International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, R. Salakhutdinov, Z. Kolter, K. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp, Eds., vol. 235. PMLR, 21–27 Jul 2024, pp. 62 429–62 442. [Online]. Available: https://proceedings.mlr.press/v235/zhu24f.html
2024
Closest in time.
H. Dubey, A. Aazami, V. Gopal, B. Naderi, S. Braun, R. Cutler, A. Ju, M. Zohourian, M. Tang, M. Golestaneh et al. , “Icassp 2023 deep noise suppression challenge,” IEEE Open Journal of Signal Processing , 2024
2024
Closest in time.
X. Liu, Q. Kong, Y. Zhao, H. Liu, Y. Yuan, Y. Liu, R. Xia, Y. Wang, M. D. Plumbley, and W. Wang, “Separate anything you describe,” IEEE Transactions on Audio, Speech and Language Processing , vol. 33, pp. 458–471, 2025
2025
Closest in time.
X. Jiang, Y. A. Li, A. Nicolas Florea, C. Han, and N. Mesgarani, “Speech slytherin: Examining the performance and efficiency of mamba for speech separation, recognition, and synthesis,” in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2025, pp. 1–5
2025
Closest in time.