Fetching the paper…
Reading the bibliography…
With the advent of the big data and large language model era, zero-shot personalized rapid customization has emerged as a significant trend.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Deep voice: Real-time neural text-to-speech
Sercan Ö Arık, Mike Chrzanowski, Adam Coates, Gregory Diamos, Andrew Gibiansky, Yongguo Kang, Xian Li, John Miller, Andrew Ng, Jonathan Raiman, et al · 2017
Earlier work this paper cites.
Tacotron: Towards end-to-end speech synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al · 2017
Earlier work this paper cites.
Natural TTS synthesis by conditioning wavenet on MEL spectrogram predictions
Jonathan Shen, Ruoming Pang, Ron J. Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, RJ-Skerrv Ryan, Rif A. Saurous, Yannis Agiomyrgiannakis, and Yonghui Wu · 2018
Earlier work this paper cites.
Low bit-rate speech coding with vq-vae and a wavenet decoder
Cristina Gârbacea, Aäron van den Oord, Yazhe Li, Felicia SC Lim, Alejandro Luebs, Oriol Vinyals, and Thomas C Walters · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2020
Earlier work this paper cites.
Conformer: Convolution-augmented transformer for speech recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al · 2020
Earlier work this paper cites.
Simbert: Integrating retrieval and generation into bert
Jianlin Su · 2020
Earlier work this paper cites.
Libri-light: A benchmark for asr with limited or no supervision
Jacob Kahn, Morgane Riviere, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, Julien Karadayi, Vitaliy Liptchinsky, Ronan Collobert, Christian Fuegen, et al · 2020
Earlier work this paper cites.
Fastspeech 2: Fast and high-quality end-to-end text to speech
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2020
Earlier work this paper cites.
Soundstream: An end-to-end neural audio codec
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models, 2021
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech
Jaehyeon Kim, Jungil Kong, and Juhee Son · 2021
Earlier work this paper cites.
High fidelity neural audio compression
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Earlier work this paper cites.
Flow matching for generative modeling
Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le · 2022
Earlier work this paper cites.
Disentangling content and fine-grained prosody information via hybrid asr bottleneck features for voice conversion
Xintao Zhao, Feng Liu, Changhe Song, Zhiyong Wu, Shiyin Kang, Deyi Tuo, and Helen Meng · 2022
Earlier work this paper cites.
Training robust zero-shot voice conversion models with self-supervised features
Trung Dang, Dung Tran, Peter Chin, and Kazuhito Koishida · 2022
Earlier work this paper cites.
Squeezeformer: An efficient transformer for automatic speech recognition
Sehoon Kim, Amir Gholami, Albert Shaw, Nicholas Lee, Karttikeya Mangalam, Jitendra Malik, Michael W Mahoney, and Kurt Keutzer · 2022
Cited alongside, same era.
Lmec: Learnable multiplicative absolute position embedding based conformer for speech recognition
Yuguang Yang, Yu Pan, Jingjing Yin, and Heng Lu · 2022
Cited alongside, same era.
Revisiting over-smoothness in text to speech
Yi Ren, Xu Tan, Tao Qin, Zhou Zhao, and Tie-Yan Liu · 2022
Cited alongside, same era.
Utmos: Utokyo-sarulab system for voicemos challenge 2022
Takaaki Saeki, Detai Xin, Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, and Hiroshi Saruwatari · 2022
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Cam++: A fast and efficient network for speaker verification using context-aware masking
Hui Wang, Siqi Zheng, Yafeng Chen, Luyao Cheng, and Qian Chen · 2023
Later among the works it cites.
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao · 2023
Later among the works it cites.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica · 2023
Later among the works it cites.
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han · 2023
Later among the works it cites.
Qwen2-audio technical report, 2024
Yunfei Chu, Jin Xu, Qian Yang, Haojie Wei, Xipin Wei, Zhifang Guo, Yichong Leng, Yuanjun Lv, Jinzheng He, Junyang Lin, Chang Zhou, and Jingren Zhou · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · 2022
Cited alongside, same era.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh · 2022
Cited alongside, same era.
Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition
Binbin Zhang, Hang Lv, Pengcheng Guo, Qijie Shao, Chao Yang, Lei Xie, Xin Xu, Hui Bu, Xiaoyu Chen, Chenchen Zeng, et al · 2022
Cited alongside, same era.
Conditional flow matching: Simulation-free dynamic optimal transport
Alexander Tong, Nikolay Malkin, Guillaume Huguet, Yanlei Zhang, Jarrid Rector-Brooks, Kilian Fatras, Guy Wolf, and Yoshua Bengio · 2023
Cited alongside, same era.
Neural codec language models are zero-shot text to speech synthesizers
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al · 2023
Cited alongside, same era.
Better speech synthesis through scaling
James Betker · 2023
Cited alongside, same era.
Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers
Kai Shen, Zeqian Ju, Xu Tan, Yanqing Liu, Yichong Leng, Lei He, Tao Qin, Sheng Zhao, and Jiang Bian · 2023
Cited alongside, same era.
Lm-vc: Zero-shot voice conversion via speech generation based on language models
Zhichao Wang, Yuanzhe Chen, Lei Xie, Qiao Tian, and Yuping Wang · 2023
Cited alongside, same era.
Closest in time.
Chatglm: A family of large language models from glm-130b to glm-4 all tools, 2024
Team GLM, :, Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Diego Rojas, Guanyu Feng, Hanlin Zhao, Hanyu Lai, Hao Yu, Hongning Wang, Jiadai Sun, Jiajie Zhang, Jiale Cheng, Jiayi Gui, Jie Tang, Jing Zhang, Juanzi Li, Lei Zhao, Lindong Wu, Lucen Zhong, Mingdao Liu, Minlie Huang, Peng Zhang, Qinkai Zheng, Rui Lu, Shuaiqi Duan, Shudan Zhang, Shulin Cao, Shuxun Yang, Weng Lam Tam, Wenyi Zhao, Xiao Liu, Xiao Xia, Xiaohan Zhang, Xiaotao Gu, Xin Lv, Xinghan Liu, Xinyi Liu, Xinyue Yang, Xixuan Song, Xunkai Zhang, Yifan An, Yifan Xu, Yilin Niu, Yuantao Yang, Yueyan Li, Yushi Bai, Yuxiao Dong, Zehan Qi, Zhaoyu Wang, Zhen Yang, Zhengxiao Du, Zhenyu Hou, and Zihan Wang · 2024
Closest in time.
Yu Pan, Lei Ma, and Jianjun Zhao · 2024
Closest in time.
Base tts: Lessons from building a billion-parameter text-to-speech model on 100k hours of data
Mateusz Łajszczak, Guillermo Cámbara, Yang Li, Fatih Beyhan, Arent van Korlaar, Fan Yang, Arnaud Joly, Álvaro Martín-Cortinas, Ammar Abbas, Adam Michalski, et al · 2024
Closest in time.
Naturalspeech: End-to-end text-to-speech synthesis with human-level quality
Xu Tan, Jiawei Chen, Haohe Liu, Jian Cong, Chen Zhang, Yanqing Liu, Xi Wang, Yichong Leng, Yuanhao Yi, Lei He, et al · 2024
Closest in time.
Promptvc: Flexible stylistic voice conversion in latent space driven by natural language prompts
Jixun Yao, Yuguang Yang, Yi Lei, Ziqian Ning, Yanni Hu, Yu Pan, Jingjing Yin, Hongbin Zhou, Heng Lu, and Lei Xie · 2024
Closest in time.
https://proceedings.neurips.cc/paper_files/paper/2023/hash/a85b405ed65c6477a4fe8302b5e06ce7-Abstract-Conference.html
Direct Preference Optimization: Your Language Model is Secretly a Reward Model — proceedings.neurips.cc · 2024
Closest in time.
Kto: Model alignment as prospect theoretic optimization
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela · 2024
Closest in time.
Binary classifier optimization for large language model alignment
Seungjae Jung, Gunsoo Han, Daniel Wontae Nam, and Kyoung-Woon On · 2024
Closest in time.
https://arxiv.org/abs/2404.05600
SpeechAlign: Aligning Speech Generation to Human Preferences — arxiv.org · 2024
Closest in time.
https://arxiv.org/abs/2406.00654
Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback — arxiv.org · 2024
Closest in time.
https://arxiv.org/abs/2407.02243
Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization — arxiv.org · 2024
Closest in time.
Gemo-clap: Gender-attribute-enhanced contrastive language-audio pretraining for accurate speech emotion recognition
Yu Pan, Yanni Hu, Yuguang Yang, Wen Fei, Jixun Yao, Heng Lu, Lei Ma, and Jianjun Zhao · 2024
Closest in time.
Yu Pan, Yuguang Yang, Heng Lu, Lei Ma, and Jianjun Zhao · 2024
Closest in time.
https://arxiv.org/abs/2407.21787
Large Language Monkeys: Scaling Inference Compute with Repeated Sampling — arxiv.org · 2024
Closest in time.