Fetching the paper…
Reading the bibliography…
In this paper, we present TridentSE, a novel architecture for speech enhancement, which is capable of efficiently capturing both global information and local details.
“Speech segregation based on pitch tracking and amplitude modulation,”
Guoning Hu and DeLiang Wang, · 2001
Earlier work this paper cites.
“Binary and ratio time-frequency masks for robust speech recognition,”
Soundararajan Srinivasan, Nicoleta Roman, and DeLiang Wang, · 2006
Earlier work this paper cites.
“Evaluation of objective quality measures for speech enhancement,”
Yi Hu and Philipos C Loizou, · 2007
Earlier work this paper cites.
“The voice bank corpus: Design, collection and data analysis of a large regional accent speech database,”
Christophe Veaux, Junichi Yamagishi, and Simon King, · 2013
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks,”
Hakan Erdogan, John R Hershey, Shinji Watanabe, and Jonathan Le Roux, · 2015
Earlier work this paper cites.
“Complex ratio masking for monaural speech separation,”
Donald S Williamson, Yuxuan Wang, and DeLiang Wang, · 2015
Earlier work this paper cites.
“Investigating rnn-based speech enhancement methods for noise-robust text-to-speech,”
Cassia Valentini-Botinhao, Xin Wang, Shinji Takaki, and Junichi Yamagishi, · 2016
Earlier work this paper cites.
“A two-stage algorithm for noisy and reverberant speech enhancement,”
Yan Zhao, Zhong-Qiu Wang, and DeLiang Wang, · 2017
Earlier work this paper cites.
“SEGAN: speech enhancement generative adversarial network,”
Santiago Pascual, Antonio Bonafonte, and Joan Serrà, · 2017
Earlier work this paper cites.
“Conv-tasnet: Surpassing ideal time-frequency magnitude masking for speech separation,”
Yi Luo and Nima Mesgarani, · 2019
Cited alongside, same era.
“Metricgan: Generative adversarial networks based black-box metric scores optimization for speech enhancement,”
Szu-Wei Fu, Chien-Feng Liao, Yu Tsao, and Shou-De Lin, · 2019
Cited alongside, same era.
“Real time speech enhancement in the waveform domain,”
Alexandre Défossez, Gabriel Synnaeve, and Yossi Adi, · 2020
Cited alongside, same era.
“PHASEN: A phase-and-harmonics-aware speech enhancement network,”
Dacheng Yin, Chong Luo, Zhiwei Xiong, and Wenjun Zeng, · 2020
Cited alongside, same era.
“DCCRN: deep complex convolution recurrent network for phase-aware speech enhancement,”
Yanxin Hu, Yun Liu, Shubo Lv, Mengtao Xing, Shimin Zhang, Yihui Fu, Jian Wu, Bihong Zhang, and Lei Xie, · 2020
Cited alongside, same era.
Umut Isik, Ritwik Giri, Neerad Phansalkar, Jean-Marc Valin, Karim Helwani, and Arvindh Krishnaswamy, · 2020
Later among the works it cites.
“Se-conformer: Time-domain speech enhancement using conformer,”
Eesung Kim and Hyeji Seo, · 2021
Later among the works it cites.
“Icassp 2021 deep noise suppression challenge,”
Chandan KA Reddy, Harishchandra Dubey, Vishak Gopal, Ross Cutler, Sebastian Braun, Hannes Gamper, Robert Aichner, and Sriram Srinivasan, · 2021
Later among the works it cites.
“Interactive speech and noise modeling for speech enhancement,”
Chengyu Zheng, Xiulian Peng, Yuan Zhang, Sriram Srinivasan, and Yan Lu, · 2021
Later among the works it cites.
“Fullsubnet: A full-band and sub-band fusion model for real-time single-channel speech enhancement,”
Xiang Hao, Xiangdong Su, Radu Horaud, and Xiaofei Li, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chuanxin Tang, Chong Luo, Zhiyuan Zhao, Wenxuan Xie, and Wenjun Zeng, · 2020
Cited alongside, same era.
Chandan KA Reddy, Ebrahim Beyrami, Harishchandra Dubey, Vishak Gopal, Roger Cheng, Ross Cutler, Sergiy Matusevych, Robert Aichner, Ashkan Aazami, Sebastian Braun, et al., · 2020
Cited alongside, same era.
“Large batch optimization for deep learning: Training BERT in 76 minutes,”
Yang You, Jing Li, Sashank J. Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh, · 2020
Cited alongside, same era.
“Sudo rm-rf: Efficient networks for universal audio source separation,”
Efthymios Tzinis, Zhepei Wang, and Paris Smaragdis, · 2020
Cited alongside, same era.
“Dpt-fsnet: Dual-path transformer based full-band and sub-band fusion network for speech enhancement,”
Feng Dang, Hangting Chen, and Pengyuan Zhang, · 2022
Closest in time.
“Uformer: A unet based dilated complex & real dual-path conformer network for simultaneous speech enhancement and dereverberation,”
Yihui Fu, Yun Liu, Jingdong Li, Dawei Luo, Shubo Lv, Yukai Jv, and Lei Xie, · 2022
Closest in time.
“Dual-branch attention-in-attention transformer for single-channel speech enhancement,”
Guochen Yu, Andong Li, Chengshi Zheng, Yinuo Guo, Yutian Wang, and Hui Wang, · 2022
Closest in time.
“Cmgan: Conformer-based metric-gan for monaural speech enhancement,”
Sherif Abdulatif, Ruizhe Cao, and Bin Yang, · 2022
Closest in time.