Fetching the paper…
Reading the bibliography…
We propose MusicRL, the first music generation system finetuned from human feedback.
Rank analysis of incomplete block designs: I. the method of paired comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
A practical approach to eighteenth-century counterpoint
R. Gauldin · 1988
Earlier work this paper cites.
Minimizing ecological gaps in interface design
J. C. Thomas and W. A. Kellogg · 1989
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
The concept of ecological validity: What are its limitations and is it bad to be invalid?
D. J. Lewkowicz · 2001
Earlier work this paper cites.
Wilcoxon-signed-rank test
D. Rey and M. Neuhäuser · 2011
Earlier work this paper cites.
Cross-cultural perspectives on music and musicality
S. E. Trehub, J. Becker, and I. Morley · 2015
Earlier work this paper cites.
Sequence level training with recurrent neural networks
M. Ranzato, S. Chopra, M. Auli, and W. Zaremba · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, J. Klingner, A. Shah, M. Johnson, X. Liu, L. Kaiser, S. Gouws, Y. Kato, T. Kudo, H. Kazawa, K. Stevens, G. Kurian, N. Patil, W. Wang, C. Young, J. Smith, J. Riesa, A. Rudnick, O. Vinyals, G. Corrado, M. Hughes, and J. Dean · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Neural audio synthesis of musical notes with wavenet autoencoders
J. H. Engel, C. Resnick, A. Roberts, S. Dieleman, M. Norouzi, D. Eck, and K. Simonyan · 2017
Earlier work this paper cites.
Objective-reinforced generative adversarial networks (organ) for sequence generation models
G. L. Guimaraes, B. Sanchez-Lengeling, C. Outeiral, P. L. C. Farias, and A. Aspuru-Guzik · 2017
Earlier work this paper cites.
Sequence tutor: Conservative fine-tuning of sequence generation models with kl-control
N. Jaques, S. Gu, D. Bahdanau, J. M. Hernández-Lobato, R. E. Turner, and D. Eck · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
SING: symbol-to-instrument neural generator
A. Défossez, N. Zeghidour, N. Usunier, L. Bottou, and F. R. Bach · 2018
Earlier work this paper cites.
Bach2bach: generating music using a deep reinforcement learning approach
N. Kotecha · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
Learning to extract coherent summary via deep reinforcement learning
Y. Wu and B. Hu · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
N. Jaques, A. Ghandeharioun, J. H. Shen, C. Ferguson, A. Lapedriza, N. Jones, S. Gu, and R. Picard · 2019
Earlier work this paper cites.
Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms
K. Kilgour, M. Zuluaga, D. Roblek, and M. Sharifi · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Cited alongside, same era.
Jukebox: A generative model for music
P. Dhariwal, H. Jun, C. Payne, J. W. Kim, A. Radford, and I. Sutskever · 2020
Cited alongside, same era.
DDSP: differentiable digital signal processing
J. H. Engel, L. Hantrakul, C. Gu, and A. Roberts · 2020
Cited alongside, same era.
Rl-duet: Online music accompaniment generation using deep reinforcement learning
N. Jiang, S. Jin, Z. Duan, and C. Zhang · 2020
Cited alongside, same era.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Cited alongside, same era.
Simple and controllable music generation, 2023
J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y. Adi, and A. Défossez · 2023
Later among the works it cites.
Reward model ensembles help mitigate overoptimization
T. Coste, U. Anwar, R. Kirk, and D. Krueger · 2023
Later among the works it cites.
Vampnet: Music generation via masked acoustic token modeling, 2023
H. F. Garcia, P. Seetharaman, R. Kumar, and B. Pardo · 2023
Later among the works it cites.
Gemini: A family of highly capable multimodal models
G. Gemini Team · 2023
Later among the works it cites.
Noise2music: Text-conditioned music generation with diffusion models
Q. Huang, D. S. Park, T. Wang, T. I. Denk, A. Ly, N. Chen, Z. Zhang, Z. Zhang, J. Yu, C. Frank, et al · 2023
Later among the works it cites.
Speak, read and prompt: High-fidelity text-to-speech with minimal supervision
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Chung, Y. Zhang, W. Han, C. Chiu, J. Qin, R. Pang, and Y. Wu · 2021
Cited alongside, same era.
A generative model for creating musical rhythms with deep reinforcement learning
S. M. Karbasi, H. S. Haug, M.-K. Kvalsund, M. J. Krzyzaniak, and J. Tørresen · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Cited alongside, same era.
High fidelity neural audio compression
A. Défossez, J. Copet, G. Synnaeve, and Y. Adi · 2022
Cited alongside, same era.
Clap: Learning audio concepts from natural language supervision, 2022
B. Elizalde, S. Deshmukh, M. A. Ismail, and H. Wang · 2022
Cited alongside, same era.
Riffusion - Stable diffusion for real-time music generation, 2022
S. Forsgren and H. Martiros · 2022
Cited alongside, same era.
General-purpose, long-context autoregressive modeling with perceiver AR
C. Hawthorne, A. Jaegle, C. Cangea, S. Borgeaud, C. Nash, M. Malinowski, S. Dieleman, O. Vinyals, M. M. Botvinick, I. Simon, H. Sheahan, N. Zeghidour, J. Alayrac, J. Carreira, and J. H. Engel · 2022
Cited alongside, same era.
E. Kharitonov, D. Vincent, Z. Borsos, R. Marinier, S. Girgin, O. Pietquin, M. Sharifi, M. Tagliasacchi, and N. Zeghidour · 2023
Later among the works it cites.
A survey on deep reinforcement learning for audio-based applications
S. Latif, H. Cuayáhuitl, F. Pervez, F. Shamshad, H. S. Ali, and E. Cambria · 2023
Later among the works it cites.
Aligning text-to-image models using human feedback
K. Lee, H. Liu, M. Ryu, O. Watkins, Y. Du, C. Boutilier, P. Abbeel, M. Ghavamzadeh, and S. S. Gu · 2023
Later among the works it cites.
Audioldm: Text-to-audio generation with latent diffusion models
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. P. Mandic, W. Wang, and M. D. Plumbley · 2023
Later among the works it cites.
Gpt-4 technical report
OpenAI · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn · 2023
Later among the works it cites.
Factually consistent summarization via reinforcement learning with textual entailment feedback
P. Roit, J. Ferret, L. Shani, R. Aharoni, G. Cideron, R. Dadashi, M. Geist, S. Girgin, L. Hussenot, O. Keller, N. Momchev, S. R. Garea, P. Stanczyk, N. Vieillard, O. Bachem, G. Elidan, A. Hassidim, O. Pietquin, and I. Szpektor · 2023
Later among the works it cites.
Moûsai: Text-to-music generation with long-context latent diffusion, 2023
F. Schneider, O. Kamal, Z. Jin, and B. Schölkopf · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Later among the works it cites.
The neuroscience of music–towards ecological validity
M. Tervaniemi · 2023
Later among the works it cites.
Diffusion model alignment using direct preference optimization
B. Wallace, M. Dang, R. Rafailov, L. Zhou, A. Lou, S. Purushwalkam, S. Ermon, C. Xiong, S. Joty, and N. Naik · 2023
Later among the works it cites.
Neural codec language models are zero-shot text to speech synthesizers
C. Wang, S. Chen, Y. Wu, Z. Zhang, L. Zhou, S. Liu, Z. Chen, Y. Liu, H. Wang, J. Li, L. He, S. Zhao, and F. Wei · 2023
Later among the works it cites.
Uniaudio: An audio foundation model toward universal audio generation, 2023
D. Yang, J. Tian, X. Tan, R. Huang, S. Liu, X. Chang, J. Shi, S. Zhao, J. Bian, X. Wu, Z. Zhao, S. Watanabe, and H. Meng · 2023
Later among the works it cites.
Megabyte: Predicting million-byte sequences with multiscale transformers
L. Yu, D. Simig, C. Flaherty, A. Aghajanyan, L. Zettlemoyer, and M. Lewis · 2023
Later among the works it cites.
Stemgen: A music generation model that listens, 2024
J. D. Parker, J. Spijkervet, K. Kosta, F. Yesiler, B. Kuznetsov, J.-C. Wang, M. Avent, J. Chen, and D. Le · 2024
Closest in time.
Warm: On the benefits of weight averaged reward models, 2024
A. Ramé, N. Vieillard, L. Hussenot, R. Dadashi, G. Cideron, O. Bachem, and J. Ferret · 2024
Closest in time.