Fetching the paper…
Reading the bibliography…
Language models (LMs) have recently flourished in natural language processing and computer vision, generating high-fidelity texts or images in various tasks.
A new algorithm for data compression
Philip Gage · 1994
Earlier work this paper cites.
Verbmobil: Foundations of Speech-to-Speech Translation , by wolfgang wahlster (editor). springer, 2000. ISBN 3-540-67783-6. price £44.50 (hardback). xii+679 pages
Jason Baldridge · 2004
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli · 2004
Earlier work this paper cites.
A time delay neural network architecture for efficient modeling of long temporal contexts
Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Autoencoding beyond pixels using a learned similarity metric
Anders Boesen Lindbo Larsen, Søren Kaae Sønderby, Hugo Larochelle, and Ole Winther · 2016
Earlier work this paper cites.
Superseded - cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit
Christophe Veaux, Junichi Yamagishi, and Kirsten MacDonald · 2016
Earlier work this paper cites.
Melgan: Generative adversarial networks for conditional waveform synthesis
Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brébisson, Yoshua Bengio, and Aaron C. Courville · 2019
Earlier work this paper cites.
Libritts: A corpus derived from librispeech for text-to-speech
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J. Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification
Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck · 2020
Earlier work this paper cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae · 2020
Earlier work this paper cites.
Emotional voice conversion with cycle-consistent adversarial network, 2020
Songxiang Liu, Yuewen Cao, and Helen Meng · 2020
Earlier work this paper cites.
w2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training
Yu-An Chung, Yu Zhang, Wei Han, Chung-Cheng Chiu, James Qin, Ruoming Pang, and Yonghui Wu · 2021
Cited alongside, same era.
Wenet: Production oriented streaming and non-streaming end-to-end speech recognition toolkit
Zhuoyuan Yao, Di Wu, Xiong Wang, Binbin Zhang, Fan Yu, Chao Yang, Zhendong Peng, Xiaoyu Chen, Lei Xie, and Xin Lei · 2021
Cited alongside, same era.
Hifi++: a unified framework for neural vocoding, bandwidth extension and speech enhancement
Pavel Andreev, Aibek Alanov, Oleg Ivanov, and Dmitry P. Vetrov · 2022
Cited alongside, same era.
Yourtts: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone
Edresson Casanova, Julian Weber, Christopher Dane Shulby, Arnaldo Cândido Júnior, Eren Gölge, and Moacir A. Ponti · 2022
Cited alongside, same era.
Wavlm: Large-scale self-supervised pre-training for full stack speech processing
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, Jian Wu, Long Zhou, Shuo Ren, Yanmin Qian, Yao Qian, Jian Wu, Michael Zeng, Xiangzhan Yu, and Furu Wei · 2022
Voice conversion with just nearest neighbors
Matthew Baas, Benjamin van Niekerk, and Herman Kamper · 2023
Closest in time.
Better speech synthesis through scaling
James Betker · 2023
Closest in time.
Gigas2s: Large scale english-to-x speech-to-speech translation
Bytedance · 2023
Closest in time.
Speak, read and prompt: High-fidelity text-to-speech with minimal supervision
Eugene Kharitonov, Damien Vincent, Zalán Borsos, Raphaël Marinier, Sertan Girgin, Olivier Pietquin, Matthew Sharifi, Marco Tagliasacchi, and Neil Zeghidour · 2023
Closest in time.
Voicebox: Text-guided multilingual universal speech generation at scale
Matthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer, Leda Sari, Rashel Moritz, Mary Williamson, Vimal Manohar, Yossi Adi, Jay Mahadeokar, and Wei-Ning Hsu · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
High fidelity neural audio compression
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi · 2022
Cited alongside, same era.
Deepfilternet2: Towards real-time speech enhancement on embedded devices for full-band audio, 2022
Hendrik Schröter, Andreas K. Maier, Alberto N. Escalante-B., and Tobias Rosenkranz · 2022
Cited alongside, same era.
Privacy and utility of x-vector based speaker anonymization
Brij Mohan Lal Srivastava, Mohamed Maouche, Md. Sahidullah, Emmanuel Vincent, Aurélien Bellet, Marc Tommasi, Natalia A. Tomashenko, Xin Wang, and Junichi Yamagishi · 2022
Cited alongside, same era.
Gigast: A 10, 000-hour pseudo speech translation corpus
Rong Ye, Chengqi Zhao, Tom Ko, Chutong Meng, Tao Wang, Mingxuan Wang, and Jun Cao · 2022
Cited alongside, same era.
Deid-vc: Speaker de-identification via zero-shot pseudo voice conversion
Ruibin Yuan, Yuxuan Wu, Jacob Li, and Jaxter Kim · 2022
Cited alongside, same era.
Soundstream: An end-to-end neural audio codec
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi · 2022
Cited alongside, same era.
WENETSPEECH: A 10000+ hours multi-domain mandarin corpus for speech recognition
Binbin Zhang, Hang Lv, Pengcheng Guo, Qijie Shao, Chao Yang, Lei Xie, Xin Xu, Hui Bu, Xiaoyu Chen, Chenchen Zeng, Di Wu, and Zhendong Peng · 2022
Cited alongside, same era.
Closest in time.
Translatotron 3: Speech to speech translation with monolingual data
Eliya Nachmani, Alon Levkovitch, Yifan Ding, Chulayuth Asawaroengchai, Heiga Zen, and Michelle Tadmor Ramanovich · 2023
Closest in time.
Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers
Kai Shen, Zeqian Ju, Xu Tan, Yanqing Liu, Yichong Leng, Lei He, Tao Qin, Sheng Zhao, and Jiang Bian · 2023
Closest in time.
Styles2st: Zero-shot style transfer for direct speech-to-speech translation
Kun Song, Yi Ren, Yi Lei, Chunfeng Wang, Kun Wei, Lei Xie, Xiang Yin, and Zejun Ma · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Closest in time.
Audiodec: An open-source streaming high-fidelity neural audio codec
Yi-Chiao Wu, Israel D. Gebru, Dejan Marković, and Alexander Richard · 2023
Closest in time.
Speak foreign languages with your own voice: Cross-lingual neural codec language modeling
Ziqiang Zhang, Long Zhou, Chengyi Wang, Sanyuan Chen, Yu Wu, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, Lei He, Sheng Zhao, and Furu Wei · 2023
Closest in time.
Multi-speaker expressive speech synthesis via multiple factors decoupling
Xinfa Zhu, Yi Lei, Kun Song, Yongmao Zhang, Tao Li, and Lei Xie · 2023
Closest in time.