Fetching the paper…
Reading the bibliography…
We introduce Seed-Music, a suite of music generation systems capable of producing high-quality music with fine-grained style control.
Gansynth: Adversarial neural audio synthesis, 2019
Jesse Engel, Kumar Krishna Agrawal, Shuo Chen, Ishaan Gulrajani, Chris Donahue, and Adam Roberts · 1902
Earlier work this paper cites.
Counterpoint by convolution, 2019
Cheng-Zhi Anna Huang, Tim Cooijmans, Adam Roberts, Aaron Courville, and Douglas Eck · 1903
Earlier work this paper cites.
Non-autoregressive neural text-to-speech, 2020
Kainan Peng, Wei Ping, Zhao Song, and Kexin Zhao · 1905
Earlier work this paper cites.
Experiments in musical intelligence (emi): Non-linear linguistic-based composition
David Cope · 1989
Earlier work this paper cites.
Spectral modeling synthesis: A sound analysis/synthesis system based on a deterministic plus stochastic decomposition
Xavier Serra and Julius Smith · 1990
Earlier work this paper cites.
Singing voice synthesis: History, current work, and future directions
Perry R Cook · 1996
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Ddsp: Differentiable digital signal processing, 2020
Jesse Engel, Lamtharn Hantrakul, Chenjie Gu, and Adam Roberts · 2001
Earlier work this paper cites.
Jukebox: A generative model for music, 2020
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever · 2005
Earlier work this paper cites.
Denoising diffusion probabilistic models, 2020
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2006
Earlier work this paper cites.
The music producer’s handbook
Bobby Owsinski · 2010
Earlier work this paper cites.
Denoising diffusion implicit models, 2022
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2010
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations, 2021
Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole · 2011
Earlier work this paper cites.
Mixing Secrets
Mike Senior · 2012
Earlier work this paper cites.
Music transcription modelling and composition using deep learning, 2016
Bob L. Sturm, João Felipe Santos, Oded Ben-Tal, and Iryna Korshunova · 2016
Earlier work this paper cites.
Singing voice synthesis based on deep neural networks
Masanari Nishimura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, and Keiichi Tokuda · 2016
Earlier work this paper cites.
Performance rnn: Generating music with expressive timing and dynamics
Ian Simon and Sageev Oore · 2017
Earlier work this paper cites.
Neural audio synthesis of musical notes with wavenet autoencoders, 2017
Jesse Engel, Cinjon Resnick, Adam Roberts, Sander Dieleman, Douglas Eck, Karen Simonyan, and Mohammad Norouzi · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Noam Shazeer, Ian Simon, Curtis Hawthorne, Andrew M. Dai, Matthew D. Hoffman, Monica Dinculescu, and Douglas Eck · 2018
Earlier work this paper cites.
Efficient neural audio synthesis, 2018
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron van den Oord, Sander Dieleman, and Koray Kavukcuoglu · 2018
Earlier work this paper cites.
Minimum word error rate training for attention-based sequence-to-sequence models
Rohit Prabhavalkar, Tara N Sainath, Yonghui Wu, Patrick Nguyen, Zhifeng Chen, Chung-Cheng Chiu, and Anjuli Kannan · 2018
Earlier work this paper cites.
Neural voice cloning with a few samples
Sercan Arik, Jitong Chen, Kainan Peng, Wei Ping, and Yanqi Zhou · 2018
Earlier work this paper cites.
Universality and diversity in human song
Samuel Mehr, Manvir Singh, Dean Knox, Daniel Ketter, Daniel Pickens-Jones, S Atwood, Christopher Lucas, Nori Jacoby, Alena Egner, Erin Hopkins, Rhea Howard, Joshua Hartshorne, Mariela Jennings, Jan Simson, Constance Bainbridge, Steven Pinker, Timothy O’Donnell, Max Krasnow, and Luke Glowacki · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners, 2019
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Fast and flexible neural audio synthesis, 2019
Lamtharn (Hanoi) Hantrakul, Jesse Engel, Adam Roberts, and Chenjie Gu · 2019
Earlier work this paper cites.
Enabling factorized piano music modeling and generation with the MAESTRO dataset
Curtis Hawthorne, Andriy Stasyuk, Adam Roberts, Ian Simon, Cheng-Zhi Anna Huang, Sander Dieleman, Erich Elsen, Jesse Engel, and Douglas Eck · 2019
Earlier work this paper cites.
Singing voice synthesis using deep autoregressive neural networks for acoustic modeling
Yuan-Hao Yi, Yang Ai, Zhen-Hua Ling, and Li-Rong Dai · 2019
Earlier work this paper cites.
Rethinking lossy compression: The rate-distortion-perception tradeoff
Yochai Blau and Tomer Michaeli · 2019
Earlier work this paper cites.
Muspy: A toolkit for symbolic music generation
Hao-Wen Dong, Ke Chen, Julian McAuley, and Taylor Berg-Kirkpatrick · 2020
Earlier work this paper cites.
Xiaoicesing: A high-quality and integrated singing voice synthesis system
Peiling Lu, Jie Wu, Jian Luan, Xu Tan, and Li Zhou · 2020
Earlier work this paper cites.
Pop music transformer: Beat-based modeling and generation of expressive pop piano compositions
Yu-Siang Huang and Yi-Hsuan Yang · 2020
Cited alongside, same era.
Dadagp: A dataset of tokenized guitarpro songs for sequence models
Pedro Sarmento, Adarsh Kumar, CJ Carr, Zack Zukowski, Mathieu Barthet, and Yi-Hsuan Yang · 2021
Cited alongside, same era.
LiteSing: Towards fast, lightweight and expressive singing voice synthesis
Xiaobin Zhuang, Tao Jiang, Szu-Yu Chou, Bin Wu, Peng Hu, and Simon Lui · 2021
Cited alongside, same era.
Soundstream: An end-to-end neural audio codec
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi · 2021
Cited alongside, same era.
Yu-An Chung, Yu Zhang, Wei Han, Chung-Cheng Chiu, James Qin, Ruoming Pang, and Yonghui Wu · 2021
Music controlnet: Multiple time-varying controls for music generation, 2023
Shih-Lun Wu, Chris Donahue, Shinji Watanabe, and Nicholas J. Bryan · 2023
Later among the works it cites.
Dual diffusion implicit bridges for image-to-image translation, 2023
Xuan Su, Jiaming Song, Chenlin Meng, and Stefano Ermon · 2023
Later among the works it cites.
Imagic: Text-based real image editing with diffusion models, 2023
Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani · 2023
Later among the works it cites.
A comparative study of voice conversion models with large-scale speech and singing data: The t13 systems for the singing voice conversion challenge 2023
Ryuichi Yamamoto, Reo Yoneyama, Lester Phillip Violeta, Wen-Chin Huang, and Tomoki Toda · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sdedit: Guided image synthesis and editing with stochastic differential equations
Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon · 2021
Cited alongside, same era.
Jian Cong, Shan Yang, Lei Xie, and Dan Su · 2021
Cited alongside, same era.
Basis-MelGAN: Efficient neural vocoder based on audio decomposition
Zhengxi Liu and Yanmin Qian · 2021
Cited alongside, same era.
Telemelody: Lyric-to-melody generation with a template-based two-stage method
Zeqian Ju, Peiling Lu, Xu Tan, Rui Wang, Chen Zhang, Songruoyao Wu, Kejun Zhang, Xiangyang Li, Tao Qin, and Tie-Yan Liu · 2021
Cited alongside, same era.
SpecTNT: A time-frequency transformer for music audio
Wei-Tsung Lu, Ju-Chiang Wang, Minz Won, Keunwoo Choi, and Xuchen Song · 2021
Cited alongside, same era.
Midi-ddsp: Detailed control of musical performance via hierarchical modeling, 2022
Yusong Wu, Ethan Manilow, Yi Deng, Rigel Swavely, Kyle Kastner, Tim Cooijmans, Aaron Courville, Cheng-Zhi Anna Huang, and Jesse Engel · 2022
Cited alongside, same era.
Mulan: A joint embedding of music audio and natural language
Qingqing Huang, Aren Jansen, Joonseok Lee, Ravi Ganti, Judith Yue Li, and Daniel PW Ellis · 2022
Cited alongside, same era.
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar · 2023
Later among the works it cites.
Hifi-codec: Group-residual vector quantization for high fidelity audio codec, 2023
Dongchao Yang, Songxiang Liu, Rongjie Huang, Jinchuan Tian, Chao Weng, and Yuexian Zou · 2023
Later among the works it cites.
Audiobox: Unified audio generation with natural language prompts, 2023
Apoorv Vyas, Bowen Shi, Matthew Le, Andros Tjandra, Yi-Chiao Wu, Baishan Guo, Jiemin Zhang, Xinyue Zhang, Robert Adkins, William Ngan, Jeff Wang, Ivan Cruz, Bapi Akula, Akinniyi Akinyemi, Brian Ellis, Rashel Moritz, Yael Yungster, Alice Rakotoarison, Liang Tan, Chris Summers, Carleigh Wood, Joshua Lane, Mary Williamson, and Wei-Ning Hsu · 2023
Later among the works it cites.
xval: A continuous number encoding for large language models
Siavash Golkar, Mariel Pettee, Michael Eickenberg, Alberto Bietti, Miles Cranmer, Geraud Krawezik, Francois Lanusse, Michael McCabe, Ruben Ohana, Liam Parker, et al · 2023
Later among the works it cites.
Multitrack music transcription with a time-frequency perceiver
Wei-Tsung Lu, Ju-Chiang Wang, and Yun-Ning Hung · 2023
Later among the works it cites.
Controllable music production with diffusion models and guidance gradients
Mark Levy, Bruno Di Giorgi, Floris Weers, Angelos Katharopoulos, and Tom Nickson · 2023
Later among the works it cites.
Diffusion model alignment using direct preference optimization
Bram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou, Aaron Lou, Senthil Purushwalkam, Stefano Ermon, Caiming Xiong, Shafiq Joty, and Nikhil Naik · 2023
Later among the works it cites.
Stay on topic with classifier-free guidance
Guillaume Sanchez, Honglu Fan, Alexander Spangher, Elad Levi, Pawan Sasanka Ammanamanchi, and Stella Biderman · 2023
Later among the works it cites.
Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever · 2023
Later among the works it cites.
Shansong Liu, Atin Sakkeer Hussain, Chenshuo Sun, and Ying Shan · 2023
Later among the works it cites.
Anygpt: Unified multimodal llm with discrete sequence modeling, 2024
Jun Zhan, Junqi Dai, Jiasheng Ye, Yunhua Zhou, Dong Zhang, Zhigeng Liu, Xin Zhang, Ruibin Yuan, Ge Zhang, Linyang Li, Hang Yan, Jie Fu, Tao Gui, Tianxiang Sun, Yugang Jiang, and Xipeng Qiu · 2024
Closest in time.
Seed-tts: A family of high-quality versatile speech generation models, 2024
Philip Anastassiou, Jiawei Chen, Jitong Chen, Yuanzhe Chen, Zhuo Chen, Ziyi Chen, Jian Cong, Lelai Deng, Chuang Ding, Lu Gao, Mingqing Gong, Peisong Huang, Qingqing Huang, Zhiying Huang, Yuanyuan Huo, Dongya Jia, Chumin Li, Feiya Li, Hui Li, Jiaxin Li, Xiaoyang Li, Xingxing Li, Lin Liu, Shouda Liu, Sichao Liu, Xudong Liu, Yuchen Liu, Zhengxi Liu, Lu Lu, Junjie Pan, Xin Wang, Yuping Wang, Yuxuan Wang, Zhen Wei, Jian Wu, Chao Yao, Yifeng Yang, Yuanhao Yi, Junteng Zhang, Qidi Zhang, Shuo Zhang, Wenjie Zhang, Yang Zhang, Zilin Zhao, Dejian Zhong, and Xiaobin Zhuang · 2024
Closest in time.
Simple and controllable music generation, 2024
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez · 2024
Closest in time.
Seed-asr: Understanding diverse speech and contexts with llm-based speech recognition, 2024
Ye Bai, Jingping Chen, Jitong Chen, Wei Chen, Zhuo Chen, Chuang Ding, Linhao Dong, Qianqian Dong, Yujiao Du, Kepan Gao, Lu Gao, Yi Guo, Minglun Han, Ting Han, Wenchao Hu, Xinying Hu, Yuxiang Hu, Deyu Hua, Lu Huang, Mingkun Huang, Youjia Huang, Jishuo Jin, Fanliu Kong, Zongwei Lan, Tianyu Li, Xiaoyang Li, Zeyang Li, Zehua Lin, Rui Liu, Shouda Liu, Lu Lu, Yizhou Lu, Jingting Ma, Shengtao Ma, Yulin Pei, Chen Shen, Tian Tan, Xiaogang Tian, Ming Tu, Bo Wang, Hao Wang, Yuping Wang, Yuxuan Wang, Hanzhang Xia, Rui Xia, Shuangyi Xie, Hongmin Xu, Meng Yang, Bihong Zhang, Jun Zhang, Wanyi Zhang, Yang Zhang, Yawei Zhang, Yijie Zheng, and Ming Zou · 2024
Closest in time.
BASE TTS: Lessons from building a billion-parameter text-to-speech model on 100k hours of data
Mateusz Łajszczak, Guillermo Cámbara, Yang Li, Fatih Beyhan, Arent van Korlaar, Fan Yang, Arnaud Joly, Álvaro Martín-Cortinas, Ammar Abbas, Adam Michalski, et al · 2024
Closest in time.
MVoice: Multilingual unified voice generation with discrete representation at scale, 2024
Rongjie Huang, Chunlei Zhang, Yongqi Wang, Dongchao Yang, Jinchuan Tian, Luping Liu, Zhenhui Ye, Ziyue Jiang, Xuankai Chang, Jiatong Shi, CHAO WENG, Zhou Zhao, and Dong Yu · 2024
Closest in time.
Mustango: Toward controllable text-to-music generation, 2024
Jan Melechovsky, Zixun Guo, Deepanway Ghosal, Navonil Majumder, Dorien Herremans, and Soujanya Poria · 2024
Closest in time.
Audioldm 2: Learning holistic audio generation with self-supervised pretraining, 2024
Haohe Liu, Yi Yuan, Xubo Liu, Xinhao Mei, Qiuqiang Kong, Qiao Tian, Yuping Wang, Wenwu Wang, Yuxuan Wang, and Mark D. Plumbley · 2024
Closest in time.
Joint audio and symbolic conditioning for temporally controlled text-to-music generation, 2024
Or Tal, Alon Ziv, Itai Gat, Felix Kreuk, and Yossi Adi · 2024
Closest in time.
Music2latent: Consistency autoencoders for latent audio compression, 2024
Marco Pasini, Stefan Lattner, and George Fazekas · 2024
Closest in time.
Wavtokenizer: an efficient acoustic discrete codec tokenizer for audio language modeling, 2024
Shengpeng Ji, Ziyue Jiang, Xize Cheng, Yifu Chen, Minghui Fang, Jialong Zuo, Qian Yang, Ruiqi Li, Ziang Zhang, Xiaoda Yang, Rongjie Huang, Yidi Jiang, Qian Chen, Siqi Zheng, Wen Wang, and Zhou Zhao · 2024
Closest in time.
Voicebox: Text-guided multilingual universal speech generation at scale
Matthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer, Leda Sari, Rashel Moritz, Mary Williamson, Vimal Manohar, Yossi Adi, Jay Mahadeokar, et al · 2024
Closest in time.
High-fidelity audio compression with improved RVQGAN
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Closest in time.
MusicRL: Aligning music generation to human preferences
Geoffrey Cideron, Sertan Girgin, Mauro Verzetti, Damien Vincent, Matej Kastelic, Zalán Borsos, Brian McWilliams, Victor Ungureanu, Olivier Bachem, Olivier Pietquin, et al · 2024
Closest in time.
SpeechAlign: Aligning speech generation to human preferences
Dong Zhang, Zhaowei Li, Shimin Li, Xin Zhang, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu · 2024
Closest in time.
Back to basics: Revisiting REINFORCE style optimization for learning from human feedback in LLMs
Arash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee, Julia Kreutzer, Ahmet Üstün, and Sara Hooker · 2024
Closest in time.