Fetching the paper…
Reading the bibliography…
Language models (LMs) have shown superior performances in various speech generation tasks recently, demonstrating their powerful ability for semantic context modeling.
“Some methods for classification and analysis of multivariate observations,”
James MacQueen et al., · 1967
Earlier work this paper cites.
“Evaluation of objective quality measures for speech enhancement,”
Yi Hu and Philipos C. Loizou, · 2008
Earlier work this paper cites.
“The voice bank corpus: Design, collection and data analysis of a large regional accent speech database,”
Christophe Veaux, Junichi Yamagishi, and Simon King, · 2013
Earlier work this paper cites.
“The diverse environments multi-channel acoustic noise database (demand): A database of multichannel environmental noise recordings,”
Joachim Thiemann, Nobutaka Ito, and Emmanuel Vincent, · 2013
Earlier work this paper cites.
“SEGAN: Speech Enhancement Generative Adversarial Network,”
Santiago Pascual et al., · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“A study on data augmentation of reverberant speech for robust speech recognition,”
Tom Ko, Vijayaditya Peddinti, Daniel Povey, Michael L Seltzer, and Sanjeev Khudanpur, · 2017
Earlier work this paper cites.
“Improving language understanding by generative pre-training,”
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al., · 2018
Earlier work this paper cites.
“Bridging the gap between monaural speech enhancement and recognition with distortion-independent acoustic modeling,”
Peidong Wang, Ke Tan, et al., · 2019
Earlier work this paper cites.
“Language models are unsupervised multitask learners,”
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al., · 2019
Earlier work this paper cites.
“Wham!: Extending speech separation to noisy environments,”
Gordon Wichern, Joe Antognini, Michael Flynn, Licheng Richard Zhu, Emmett McQuinn, Dwight Crow, Ethan Manilow, and Jonathan Le Roux, · 2019
Earlier work this paper cites.
“Conv-tasnet: Surpassing ideal time-frequency magnitude masking for speech separation,”
Yi Luo and Nima Mesgarani, · 2019
Earlier work this paper cites.
“A flow-based deep latent variable model for speech spectrogram modeling and enhancement,”
Nugraha et al., · 2020
Earlier work this paper cites.
“Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,”
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, · 2020
Earlier work this paper cites.
“Language models are few-shot learners,”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al., · 2020
Cited alongside, same era.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al., · 2020
Cited alongside, same era.
“Librimix: An open-source dataset for generalizable speech separation,” 2020
Joris Cosentino, Manuel Pariente, Samuele Cornell, Antoine Deleforge, and Emmanuel Vincent, · 2020
Cited alongside, same era.
“Nu-gan: High resolution neural upsampling with gan,” 2020
Rithesh Kumar, Kundan Kumar, Vicki Anand, Yoshua Bengio, and Aaron Courville, · 2020
Cited alongside, same era.
“ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification,”
Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck, · 2020
Cited alongside, same era.
“Conditional diffusion probabilistic model for speech enhancement,”
Yen-Ju Lu, Zhong-Qiu Wang, Shinji Watanabe, Alexander Richard, Cheng Yu, and Yu Tsao, · 2022
Later among the works it cites.
“Speech Enhancement with Score-Based Generative Models in the Complex STFT Domain,”
Simon Welker, Julius Richter, and Timo Gerkmann, · 2022
Later among the works it cites.
“Wavlm: Large-scale self-supervised pre-training for full stack speech processing,”
Sanyuan Chen et al., · 2022
Later among the works it cites.
“Storm: A diffusion-based stochastic regeneration model for speech enhancement and dereverberation,”
Jean-Marie Lemercier, Julius Richter, Simon Welker, and Timo Gerkmann, · 2023
Closest in time.
“Exploring wavlm on speech enhancement,”
Hyungchan Song, Sanyuan Chen, Zhuo Chen, Yu Wu, Takuya Yoshioka, Min Tang, Jong Won Shin, and Shujie Liu, · 2023
Closest in time.
“Neural codec language models are zero-shot text to speech synthesizers,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Real Time Speech Enhancement in the Waveform Domain,”
Alexandre Défossez, Gabriel Synnaeve, and Yossi Adi, · 2020
Cited alongside, same era.
“Variational autoencoder for speech enhancement with a noise-aware encoder,”
Huajian Fang, Guillaume Carbajal, Stefan Wermter, and Timo Gerkmann, · 2021
Cited alongside, same era.
“SUPERB: Speech Processing Universal PERformance Benchmark,”
Shu wen Yang et al., · 2021
Cited alongside, same era.
“On generative spoken language modeling from raw audio,”
Kushal Lakhotia, Eugene Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Abdelrahman Mohamed, et al., · 2021
Cited alongside, same era.
“Wenet: Production oriented streaming and non-streaming end-to-end speech recognition toolkit,”
Zhuoyuan Yao et al., · 2021
Cited alongside, same era.
“GigaSpeech: An Evolving, Multi-Domain ASR Corpus with 10,000 Hours of Transcribed Audio,”
Guoguo Chen et al., · 2021
Cited alongside, same era.
“INTERSPEECH 2021 Deep Noise Suppression Challenge,”
Chandan K.A. Reddy et al., · 2021
Cited alongside, same era.
Chengyi Wang et al., · 2023
Closest in time.
“Speechgen: Unlocking the generative power of speech language models with prompts,”
Haibin Wu, Kai-Wei Chang, Yuan-Kuei Wu, and Hung-yi Lee, · 2023
Closest in time.
“Speechx: Neural codec language model as a versatile speech transformer,”
Xiaofei Wang et al., · 2023
Closest in time.
Hakan Erdogan et al., · 2023
Closest in time.
“Voicebox: Text-guided multilingual universal speech generation at scale,”
Matthew Le et al., · 2023
Closest in time.
“Vec-tok speech: Speech vectorization and tokenization for neural speech generation,”
Xinfa Zhu, Yuanjun Lv, Yi Lei, Tao Li, Wendi He, Hongbin Zhou, and Lei Xie, · 2023
Closest in time.
“Audiolm: a language modeling approach to audio generation,”
Zalán Borsos et al., · 2023
Closest in time.
“Inter-subnet: Speech enhancement with subband interaction,”
Jun Chen et al., · 2023
Closest in time.