Fetching the paper…
Reading the bibliography…
We propose SE-Bridge, a novel method for speech enhancement (SE).
J. Garofolo, D. Graff, D. Paul, and D. Pallett, “Csr-i (wsj0) complete ldc93s6a,” in Web Download. Philadelphia: Linguistic Data Consortium , vol. 83, 1993
1993
Earlier work this paper cites.
C. Veaux, J. Yamagishi, and S. King, “The voice bank corpus: Design, collection and data analysis of a large regional accent speech database,” in 2013 International Conference Oriental COCOSDA held jointly with 2013 Conference on Asian Spoken Language Research and Evaluation (O-COCOSDA/CASLRE), Gurgaon, India, November 25-27, 2013 . IEEE, 2013, pp. 1–4. [Online]. Available: https://doi.org/10.1109/ICSDA.2013.6709856
2013
Earlier work this paper cites.
B. Series, “Method for the subjective assessment of intermediate quality level of audio systems,” International Telecommunication Union Radiocommunication Assembly , 2014
2014
Earlier work this paper cites.
J. Barker, R. Marxer, E. Vincent, and S. Watanabe, “The third ‘chime’speech separation and recognition challenge: Dataset, task and baselines,” in IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) . IEEE, 2015, pp. 504–511
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An asr corpus based on public domain audio books,” in IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
S. Pascual, A. Bonafonte, and J. Serrà, “SEGAN: speech enhancement generative adversarial network,” in Interspeech 2017, 18th Annual Conference of the International Speech Communication Association, Stockholm, Sweden, August 20-24, 2017 , F. Lacerda, Ed. ISCA, 2017, pp. 3642–3646. [Online]. Available: http://www.isca-speech.org/archive/Interspeech_2017/abstracts/1428.html
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Łukasz Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Inf. Proc. Systems (NeurIPS) , 2017
2017
Earlier work this paper cites.
A. Nagrani, J. S. Chung, and A. Zisserman, “Voxceleb: A large-scale speaker identification dataset,” in Proc. Interspeech 2017 , 2017, pp. 2616–2620. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2017-950
2017
Earlier work this paper cites.
M. H. Soni, N. Shah, and H. A. Patil, “Time-frequency masking-based speech enhancement using generative adversarial network,” in IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , 2018, pp. 5039–5043
2018
Earlier work this paper cites.
S. Leglaive, L. Girin, and R. Horaud, “A variance modeling framework based on variational autoencoders for speech enhancement,” in IEEE Int. Workshop on Machine Learning for Signal Proc. (MLSP) , 2018, pp. 1–6
2018
Earlier work this paper cites.
Y. Bando, M. Mimura, K. Itoyama, K. Yoshii, and T. Kawahara, “Statistical speech enhancement based on probabilistic integration of variational autoencoder and non-negative matrix factorization,” in IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , 2018, pp. 716–720
2018
Earlier work this paper cites.
D. Baby and S. Verhulst, “SERGAN: Speech enhancement using relativistic generative adversarial networks with gradient penalty,” in IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , 2019, pp. 106–110
2019
Earlier work this paper cites.
S. Leglaive, L. Girin, and R. Horaud, “Semi-supervised multichannel speech enhancement with variational autoencoders and non-negative matrix factorization,” in IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , 2019, pp. 101–105
2019
Cited alongside, same era.
S. Leglaive, U. Şimşekli, A. Liutkus, L. Girin, and R. Horaud, “Speech enhancement with variational autoencoders and alpha-stable distributions,” in IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , 2019, pp. 541–545
2019
Cited alongside, same era.
S. S¨arkk¨a and A. Solin, “Applied stochastic differential equations,” in Cambridge University Press , 2019
2019
Cited alongside, same era.
Y. Luo and N. Mesgarani, “Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation,” IEEE/ACM Trans. Audio, speech, Lang. Process. , vol. 27, no. 8, pp. 1256–1266, 2019
2019
Cited alongside, same era.
M. Sadeghi and X. Alameda-Pineda, “Mixture of inference networks for vae-based audio-visual speech enhancement,” IEEE Transactions on Signal Processing , vol. 69, pp. 1899–1909, 2021
2021
Later among the works it cites.
M. Strauss and B. Edler, “A flow-based neural network for time domain speech enhancement,” in IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , 2021, pp. 5754–5758
2021
Later among the works it cites.
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” in 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenReview.net, 2021. [Online]. Available: https://openreview.net/forum?id=PxTIG12RRHS
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Liu, K. Gong, X. Liang, and Z. Chen, “CPGAN: Context pyramid generative adversarial network for speech enhancement,” in IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , 2020, pp. 6624–6628
2020
Cited alongside, same era.
S. Leglaive, X. Alameda-Pineda, L. Girin, and R. Horaud, “A recurrent variational autoencoder for speech enhancement,” in IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , 2020, pp. 371–375
2020
Cited alongside, same era.
A. A. Nugraha, K. Sekiguchi, and K. Yoshii, “A flow-based deep latent variable model for speech spectrogram modeling and enhancement,” IEEE/ACM Trans. Audio, speech, Lang. Process. , vol. 28, pp. 1104–1117, 2020
2020
Cited alongside, same era.
P. Wang, K. Tan, and D. L. Wang, “Bridging the gap between monaural speech enhancement and recognition with distortion-independent acoustic modeling,” in IEEE Trans. on Audio, Speech, and Language Proc. (TASLP) , vol. 28, 2020, pp. 39–48
2020
Cited alongside, same era.
J. Richter, G. Carbajal, and T. Gerkmann, “Speech enhancement with stochastic temporal convolutional networks,” ISCA Interspeech , pp. 4516–4520, 2020
2020
Cited alongside, same era.
B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification,” in Proc. Interspeech 2020 , 2020, pp. 3830–3834. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2020-2650
2020
Cited alongside, same era.
S. Fu, C. Yu, T. Hsieh, P. Plantinga, M. Ravanelli, X. Lu, and Y. Tsao, “MetricGAN+: an improved version of metricgan for speech enhancement,” in Interspeech 2021, 22nd Annual Conference of the International Speech Communication Association, Brno, Czechia, 30 August - 3 September 2021 , H. Hermansky, H. Cernocký, L. Burget, L. Lamel, O. Scharenborg, and P. Motlícek, Eds. ISCA, 2021, pp. 201–205. [Online]. Available: https://doi.org/10.21437/Interspeech.2021-599
2021
Cited alongside, same era.
2021
Later among the works it cites.
S. Welker, J. Richter, and T. Gerkmann, “Speech enhancement with score-based generative models in the complex STFT domain,” in ISCA Interspeech , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
X. Bie, S. Leglaive, X. Alameda-Pineda, and L. Girin, “Unsupervised speech enhancement using dynamical variational autoencoders,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 2993–3007, 2022
2022
Later among the works it cites.
Y.-J. Lu, Z.-Q. Wang, S. Watanabe, A. Richard, C. Yu, and Y. Tsao, “Conditional diffusion probabilistic model for speech enhancement,” in IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , 2022
2022
Later among the works it cites.
2023
Closest in time.