Fetching the paper…
Reading the bibliography…
VQ-VAE, as a mainstream approach of speech tokenizer, has been troubled by ``index collapse'', where only a small number of codewords are activated in large codebooks.
“Product quantization for nearest neighbor search,”
Herve Jegou, Matthijs Douze, and Cordelia Schmid, · 2010
Earlier work this paper cites.
“Approximate nearest neighbor search by residual vector quantization,”
Yongjian Chen, Tao Guan, and Cheng Wang, · 2010
Earlier work this paper cites.
“Neural discrete representation learning,”
Aaron van den Oord, Oriol Vinyals, and koray kavukcuoglu, · 2017
Earlier work this paper cites.
“Neural discrete representation learning,”
Aäron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu, · 2017
Earlier work this paper cites.
“Decoupled weight decay regularization,”
Ilya Loshchilov and Frank Hutter, · 2017
Earlier work this paper cites.
“Low bit-rate speech coding with VQ-VAE and a WaveNet decoder,”
Cristina Gârbacea, Aäron van den Oord, Yazhe Li, Felicia SC Lim, Alejandro Luebs, Oriol Vinyals, and Thomas C Walters, · 2019
Earlier work this paper cites.
“Language models are few-shot learners,”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al., · 2020
Earlier work this paper cites.
“Scaling laws for neural language models,”
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei, · 2020
Earlier work this paper cites.
Disong Wang, Liqun Deng, Yu Ting Yeung, Xiao Chen, Xunying Liu, and Helen Meng, · 2021
Earlier work this paper cites.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Earlier work this paper cites.
“Deep learning based assessment of synthetic speech naturalness,”
Gabriel Mittag and Sebastian Möller, · 2021
Earlier work this paper cites.
“VQTTS: High-fidelity text-to-speech synthesis with self-supervised VQ acoustic feature,”
Chenpeng Du, Yiwei Guo, Xie Chen, and K. Yu, · 2022
Cited alongside, same era.
“High fidelity neural audio compression,”
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi, · 2022
Cited alongside, same era.
“Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition,”
Binbin Zhang, Hang Lv, Pengcheng Guo, Qijie Shao, Chao Yang, Lei Xie, Xin Xu, Hui Bu, Xiaoyu Chen, Chenchen Zeng, et al., · 2022
Cited alongside, same era.
“Gpt-4 technical report,” 2023
OpenAI, · 2023
Cited alongside, same era.
“Llama 2: Open foundation and fine-tuned chat models,”
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al., · 2023
Cited alongside, same era.
“Straightening out the straight-through estimator: Overcoming optimization challenges in vector quantized networks,”
Minyoung Huh, Brian Cheung, Pulkit Agrawal, and Phillip Isola, · 2023
Later among the works it cites.
“Msmc-tts: Multi-stage multi-codebook vq-vae based neural tts,”
Haohan Guo, Fenglong Xie, Xixin Wu, Frank K Soong, and Helen MengFellow, · 2023
Later among the works it cites.
“Neural codec language models are zero-shot text to speech synthesizers,”
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al., · 2023
Later among the works it cites.
“Speechtokenizer: Unified speech tokenizer for speech large language models,”
Xin Zhang, Dong Zhang, Shimin Li, Yaqian Zhou, and Xipeng Qiu, · 2023
Later among the works it cites.
“Online clustered codebook,”
Chuanxia Zheng and Andrea Vedaldi, · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Speechgen: Unlocking the generative power of speech language models with prompts,”
Haibin Wu, Kai-Wei Chang, Yuan-Kuei Wu, and Hung-yi Lee, · 2023
Cited alongside, same era.
“Better speech synthesis through scaling,”
James Betker, · 2023
Cited alongside, same era.
“Lm-vc: Zero-shot voice conversion via speech generation based on language models,”
Zhichao Wang, Yuanzhe Chen, Lei Xie, Qiao Tian, and Yuping Wang, · 2023
Cited alongside, same era.
“Towards general-purpose text-instruction-guided voice conversion,”
Chun-Yi Kuan, Chen-An Li, Tsu-Yuan Hsu, Tse-Yang Lin, Ho-Lam Chung, Kai-Wei Chang, Shuo-Yiin Chang, and Hung-yi Lee, · 2023
Cited alongside, same era.
“Speech translation with large language models: An industrial practice,”
Zhichao Huang, Rong Ye, Tom Ko, Qianqian Dong, Shanbo Cheng, Mingxuan Wang, and Hang Li, · 2023
Cited alongside, same era.
“Repcodec: A speech representation codec for speech tokenization,”
Zhichao Huang, Chutong Meng, and Tom Ko, · 2023
Cited alongside, same era.
Haohan Guo, Fenglong Xie, Jiawen Kang, Yujia Xiao, Xixin Wu, and Helen Meng, · 2023
Later among the works it cites.
“Better speech synthesis through scaling,”
James Betker, · 2023
Later among the works it cites.
“BigVGAN: A universal neural vocoder with large-scale training,”
Sang-gil Lee, Wei Ping, Boris Ginsburg, Bryan Catanzaro, and Sungroh Yoon, · 2023
Later among the works it cites.
“Finite scalar quantization: Vq-vae made simple,”
Fabian Mentzer, David Minnen, Eirikur Agustsson, and Michael Tschannen, · 2023
Later among the works it cites.
“Textually pretrained speech language models,”
Michael Hassid, Tal Remez, Tu Anh Nguyen, Itai Gat, Alexis Conneau, Felix Kreuk, Jade Copet, Alexandre Defossez, Gabriel Synnaeve, Emmanuel Dupoux, et al., · 2024
Closest in time.
“Base tts: Lessons from building a billion-parameter text-to-speech model on 100k hours of data,”
Mateusz Łajszczak, Guillermo Cámbara, Yang Li, Fatih Beyhan, Arent van Korlaar, Fan Yang, Arnaud Joly, Álvaro Martín-Cortinas, Ammar Abbas, Adam Michalski, et al., · 2024
Closest in time.