Fetching the paper…
Reading the bibliography…
Transformer and its derivatives have achieved success in diverse tasks across computer vision, natural language processing, and speech processing.
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning internal representations by error propagation, parallel distributed processing, explorations in the microstructure of cognition, ed. de rumelhart and j. mcclelland. vol. 1. 1986,” Biometrika , vol. 71, pp. 599–607, 1986
1986
Earlier work this paper cites.
H. J. Steeneken and F. W. Geurtsen, “Description of the RSG-10 noise database,” report IZF , vol. 3, p. 1988, 1988
1988
Earlier work this paper cites.
A. Acero and R. M. Stern, “Environmental robustness in automatic speech recognition,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. IEEE, 1990, pp. 849–852
1990
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
M. Harris, S. Sengupta, and J. D. Owens, “Parallel prefix sum (scan) with cuda,” GPU gems , vol. 3, no. 39, pp. 851–876, 2007
2007
Earlier work this paper cites.
R. I.-T. P. ITU, “862.2: Wideband extension to recommendation P. 862 for the assessment of wideband telephone networks and speech codecs. ITU-Telecommunication standardization sector, 2007.”
2007
Earlier work this paper cites.
Y. Hu and P. C. Loizou, “Evaluation of objective quality measures for speech enhancement,” IEEE Trans. Audio, Speech, Lang. process. , vol. 16, no. 1, pp. 229–238, 2007
2007
Earlier work this paper cites.
D. B. Dean, S. Sridharan, R. J. Vogt, and M. W. Mason, “The QUT-NOISE-TIMIT corpus for the evaluation of voice activity detection algorithms,” in Proc. Interspeech , 2010
2010
Earlier work this paper cites.
G. Hu and D. Wang, “A tandem algorithm for pitch estimation and voiced speech segregation,” IEEE/ACM Trans. Audio, speech, Lang. Process. , vol. 18, no. 8, pp. 2067–2079, 2010
2010
Earlier work this paper cites.
D.-C. Lyu, T.-P. Tan, E. S. Chng, and H. Li, “Seame: a mandarin-english code-switching speech corpus in south-east asia,” in Eleventh Annual Conference of the International Speech Communication Association , 2010
2010
Earlier work this paper cites.
H. Li, B. Ma, and K. A. Lee, “Spoken language recognition: from fundamentals to practice,” Proc. IEEE , vol. 101, no. 5, pp. 1136–1159, 2013
2013
Earlier work this paper cites.
J. Salamon, C. Jacoby, and J. P. Bello, “A dataset and taxonomy for urban sound research,” in Proc. ACM-MM , 2014, pp. 1041–1044
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2015, pp. 5206–5210
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
F. Saki and N. Kehtarnavaz, “Automatic switching between noise classification and speech enhancement for hearing aid devices,” in in Proc. EMBC , 2016, pp. 736–739
2016
Earlier work this paper cites.
F. Saki, A. Sehgal, I. Panahi, and N. Kehtarnavaz, “Smartphone-based real-time classification of noise signals using subband features and random forest classifier,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2016, pp. 2204–2208
2016
Earlier work this paper cites.
J. Jensen and C. H. Taal, “An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,” IEEE/ACM Trans. Audio, speech, Lang. Process. , vol. 24, no. 11, pp. 2009–2022, 2016
2016
Earlier work this paper cites.
C. Valentini-Botinhao, X. Wang, S. Takaki, and J. Yamagishi, “Investigating rnn-based speech enhancement methods for noise-robust text-to-speech.” in SSW , 2016, pp. 146–152
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
S. Pascual, A. Bonafonte, and J. Serrà, “SEGAN: Speech enhancement generative adversarial network,” Proc. Interspeech , pp. 3642–3646, 2017
2017
Earlier work this paper cites.
L. Dong, S. Xu, and B. Xu, “Speech-transformer: A no-recurrence sequence-to-sequence model for speech recognition,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2018, pp. 5884–5888
2018
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al. , “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. Enrique Yalta Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, and T. Ochiai, “ESPnet: End-to-end speech processing toolkit,” in Proc. Interspeech , 2018, pp. 2207–2211
2018
Earlier work this paper cites.
M. Gupta, D. Bahri, A. Cotter, and K. Canini, “Diminishing returns shape constraints for interpretability and regularization,” Advances in neural information processing systems , vol. 31, 2018
2018
Earlier work this paper cites.
J. D. M.-W. C. Kenton and L. K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of NAACL-HLT , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
Z. Dai, Z. Yang, Y. Yang, J. G. Carbonell, Q. Le, and R. Salakhutdinov, “Transformer-xl: Attentive language models beyond a fixed-length context,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019, pp. 2978–2988
2019
Earlier work this paper cites.
Z. Zeng, Y. Khassanov, V. T. Pham, H. Xu, E. S. Chng, and H. Li, “On the end-to-end solution to mandarin-english code-switching speech recognition,” in Proc. Interspeech , 2019, pp. 2165–2169
2019
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations , 2020
2020
Earlier work this paper cites.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented Transformer for Speech Recognition,” in Proc. Interspeech , 2020, pp. 5036–5040
2020
Earlier work this paper cites.
Q. Zhang, A. Nicolson, M. Wang, K. K. Paliwal, and C. Wang, “DeepMMSE: A deep learning approach to mmse-based noise power spectral density estimation,” IEEE/ACM Trans. Audio, speech, Lang. Process. , vol. 28, pp. 1404–1415, 2020
2020
Earlier work this paper cites.
A. Nicolson and K. K. Paliwal, “Masked multi-head self-attention for causal speech enhancement,” Speech Com. , vol. 125, pp. 80–96, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
H. Phan, I. V. McLoughlin, L. Pham, O. Y. Chén, P. Koch, M. De Vos, and A. Mertins, “Improving gans for speech enhancement,” IEEE Signal Processing Letters , vol. 27, pp. 1700–1704, 2020
2020
Cited alongside, same era.
T.-A. Hsieh, H.-M. Wang, X. Lu, and Y. Tsao, “Wavecrn: An efficient convolutional recurrent neural network for end-to-end speech enhancement,” IEEE Signal Processing Letters , vol. 27, pp. 2149–2153, 2020
2020
Cited alongside, same era.
Q. Zhang, X. Qian, Z. Ni, A. Nicolson, E. Ambikairajah, and H. Li, “A time-frequency attention module for neural speech enhancement,” IEEE/ACM Trans. Audio, speech, Lang. Process. , vol. 31, pp. 462–475, 2023
2023
Later among the works it cites.
Q. Zhang, X. Qian, Z. Ni, A. Nicolson, E. Ambikairajah, and H. Li, “A time-frequency attention module for neural speech enhancement,” IEEE/ACM Trans. Audio, speech, Lang. Process. , vol. 31, pp. 462–475, 2023
2023
Later among the works it cites.
Q. Zhang, H. Zhu, Q. Song, X. Qian, Z. Ni, and H. Li, “Ripple sparse self-attention for monaural speech enhancement,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2023, pp. 1–5
2023
Later among the works it cites.
J. Richter, S. Welker, J.-M. Lemercier, B. Lay, and T. Gerkmann, “Speech enhancement and dereverberation with diffusion-based generative models,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 31, pp. 2351–2364, 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Koizumi, K. Yatabe, M. Delcroix, Y. Masuyama, and D. Takeuchi, “Speech enhancement using self-adaptation and multi-head self-attention,” in Proc. ICASSP . IEEE, 2020, pp. 181–185
2020
Cited alongside, same era.
J. Su, Z. Jin, and A. Finkelstein, “Hifi-gan: High-fidelity denoising and dereverberation based on speech deep features in adversarial networks,” in Proc. INTERSPEECH , 2020
2020
Cited alongside, same era.
D. Yin, C. Luo, Z. Xiong, and W. Zeng, “Phasen: A phase-and-harmonics-aware speech enhancement network,” in Proc. AAAI , 2020, pp. 9458–9465
2020
Cited alongside, same era.
Y. Hu, Y. Liu, S. Lv, M. Xing, S. Zhang, Y. Fu, J. Wu, B. Zhang, and L. Xie, “Dccrn: Deep complex convolution recurrent network for phase-aware speech enhancement,” Proc. INTERSPEECH , pp. 2472–2476, 2020
2020
Cited alongside, same era.
B. J. Borgström and M. S. Brandstein, “Speech enhancement via attention masking network (seamnet): An end-to-end system for joint suppression of noise and reverberation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 515–526, 2020
2020
Cited alongside, same era.
J. Kim, M. El-Khamy, and J. Lee, “T-GSA: Transformer with gaussian-weighted self-attention for speech enhancement,” in Proc. ICASSP , 2020, pp. 6649–6653
2020
Cited alongside, same era.
A. Defossez, G. Synnaeve, and Y. Adi, “Real time speech enhancement in the waveform domain,” in Proc. Interspeech , 2020
2020
Cited alongside, same era.
P. Koehn, Neural machine translation . Cambridge University Press, 2020
2020
Cited alongside, same era.
J.-M. Lemercier, J. Richter, S. Welker, and T. Gerkmann, “Storm: A diffusion-based stochastic regeneration model for speech enhancement and dereverberation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2023
2023
Later among the works it cites.
Y.-X. Lu, Y. Ai, and Z.-H. Ling, “MP-SENet: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase Spectra,” in Proc. INTERSPEECH 2023 , 2023, pp. 3834–3838
2023
Later among the works it cites.
2023
Later among the works it cites.
H. Liu, H. Xu, L. P. Garcia, A. W. H. Khong, Y. He, and S. Khudanpur, “Reducing language confusion for code-switching speech recognition with token-level language diarization,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2023, pp. 1–5
2023
Later among the works it cites.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
D. Han, Z. Wang, Z. Xia, Y. Han, Y. Pu, C. Ge, J. Song, S. Song, B. Zheng, and G. Huang, “Demystify mamba in vision: A linear attention perspective,” in Proc. NeurIPS , 2024
2024
Closest in time.
S. Bhati, Y. Gong, L. Karlinsky, H. Kuehne, R. Feris, and J. Glass, “Dass: Distilled audio state space models are stronger and more duration-scalable learners,” in 2024 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2024, pp. 1015–1022
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
R. Chao, W.-H. Cheng, M. L. Quatra, S. M. Siniscalchi, C.-H. H. Yang, S.-W. Fu, and Y. Tsao, “An investigation of incorporating mamba for speech enhancement,” in Proc. IEEE Spoken Language Technology Workshop , 2024, pp. 302–308
2024
Closest in time.
K. Miyazaki, Y. Masuyama, and M. Murata, “Exploring the capability of mamba in speech applications,” in Proc. Interspeech 2024 , 2024, pp. 237–241
2024
Closest in time.
2024
Closest in time.
Q. Zhang, M. Ge, H. Zhu, E. Ambikairajah, Q. Song, Z. Ni, and H. Li, “An empirical study on the impact of positional encoding in transformer-based monaural speech enhancement,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2024, pp. 1001–1005
2024
Closest in time.
2024
Closest in time.
S. Abdulatif, R. Cao, and B. Yang, “Cmgan: Conformer-based metric-gan for monaural speech enhancement,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2024
2024
Closest in time.
H. Liu, L. P. Garcia, X. Zhang, A. W. Khong, and S. Khudanpur, “Enhancing code-switching speech recognition with interactive language biases,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. IEEE, 2024, pp. 10 886–10 890
2024
Closest in time.
C. Manning, “Lecture 8: Transformers,” CS224N: Natural Language Processing with Deep Learning, Stanford University, 2024, accessed: 2024-05-14, Slide 32. [Online]. Available: https://web.stanford.edu/class/cs224n/slides/cs224n-spr2024-lecture08-transformers.pdf
2024
Closest in time.
2024
Closest in time.
X. Gao and N. F. Chen, “Speech-mamba: Long-context speech recognition with selective state spaces models,” in 2024 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2024, pp. 1–8
2024
Closest in time.
P. Feng, Y. Wang, Y. Ni, Z. Li, W. Wu, and L. Huang, “An empirical study on normalization in mamba,” in Proc. International Conference on Learning Representations , 2025
2025
Closest in time.
M. Chen, Q. Zhang, M. Wang, X. Zhang, H. Liu, E. Ambikairaiah, and D. Chen, “Selective state space model for monaural speech enhancement,” IEEE Transactions on Consumer Electronics , 2025
2025
Closest in time.
S.-W. Fu, C.-F. Liao, Y. Tsao, and S.-D. Lin, “Metricgan: Generative adversarial networks based black-box metric scores optimization for speech enhancement,” in Proc. ICML , 2019, pp. 2031–2041
2041
Closest in time.