Fetching the paper…
Reading the bibliography…
Recent years have seen the rapid development of large generative models for text; however, much less research has explored the connection between text and another "language" of communication -- music.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
I.—COMPUTING MACHINERY AND INTELLIGENCE
Alan M. Turing. 1950 · 1950
Earlier work this paper cites.
The concept of musical syntax
Joseph P Swain. 1995 · 1995
Earlier work this paper cites.
Sonata form
James Webster. 2001 · 2001
Earlier work this paper cites.
Can music convey semantic content? a kantian approach
Jeanette Bicknell. 2002 · 2002
Earlier work this paper cites.
Jukebox: A generative model for music
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever. 2020 · 2005
Earlier work this paper cites.
Dalcroze, the body, movement and musicality
Jay A Seitz. 2005 · 2005
Earlier work this paper cites.
Digital storytelling in integrated arts education
Sheng-Kuan Chung. 2006 · 2006
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Geoffrey E Hinton and Ruslan R Salakhutdinov. 2006 · 2006
Earlier work this paper cites.
Notes , 67(4):760–765
Mark Germer. 2011 · 2011
Earlier work this paper cites.
Nicolas Boulanger-Lewandowski, Yoshua Bengio, and Pascal Vincent. 2012 · 2012
Earlier work this paper cites.
Lyrics, music, and emotions
Rada Mihalcea and Carlo Strapparava. 2012 · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling. 2014 · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015 · 2015
Earlier work this paper cites.
Proposal for a directive of the European parliament and of the council on copyright in the digital single market
European Commission. 2016 · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu. 2016 · 2016
Earlier work this paper cites.
Neural audio synthesis of musical notes with wavenet autoencoders
Jesse Engel, Cinjon Resnick, Adam Roberts, Sander Dieleman, Douglas Eck, Karen Simonyan, and Mohammad Norouzi. 2017 · 2017
Earlier work this paper cites.
SampleRNN: An unconditional end-to-end neural audio generation model
Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron C. Courville, and Yoshua Bengio. 2017 · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aäron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. 2017 · 2017
Earlier work this paper cites.
Towards a music-language mapping
Michele Berlingerio and Francesca Bonin. 2018 · 2018
Earlier work this paper cites.
The challenge of realistic music generation: Modelling raw audio at scale
Sander Dieleman, Aäron van den Oord, and Karen Simonyan. 2018 · 2018
Earlier work this paper cites.
The exception for text and data mining (tdm) in the proposed directive on copyright in the digital single market-legal aspects
Christophe Geiger, Giancarlo Frosio, and Oleksandr Bulayenko. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Gansynth: Adversarial neural audio synthesis
Jesse H. Engel, Kumar Krishna Agrawal, Shuo Chen, Ishaan Gulrajani, Chris Donahue, and Adam Roberts. 2019 · 2019
Cited alongside, same era.
Learning dense representations for entity retrieval
Daniel Gillick, Sayali Kulkarni, Larry Lansing, Alessandro Presta, Jason Baldridge, Eugene Ie, and Diego Garcia-Olano. 2019 · 2019
Cited alongside, same era.
Enabling factorized piano music modeling and generation with the MAESTRO dataset
Curtis Hawthorne, Andriy Stasyuk, Adam Roberts, Ian Simon, Cheng-Zhi Anna Huang, Sander Dieleman, Erich Elsen, Jesse H. Engel, and Douglas Eck. 2019b · 2019
Cited alongside, same era.
Fréchet audio distance: A metric for evaluating music enhancement algorithms
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi. 2019 · 2019
Cited alongside, same era.
Imagen video: High definition video generation with diffusion models
Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey A. Gritsenko, Diederik P. Kingma, Ben Poole, Mohammad Norouzi, David J. Fleet, and Tim Salimans. 2022 · 2022
Later among the works it cites.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans. 2022 · 2022
Later among the works it cites.
Commu: Dataset for combinatorial music generation
Lee Hyun, Taehyun Kim, Hyolim Kang, Minjoo Ki, Hyeonchan Hwang, Kwanho Park, Sharang Han, and Seon Joo Kim. 2022 · 2022
Later among the works it cites.
AudioGen: Textually guided audio generation
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Melgan: Generative adversarial networks for conditional waveform synthesis
Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brébisson, Yoshua Bengio, and Aaron C. Courville. 2019 · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Cited alongside, same era.
Panns: Large-scale pretrained audio neural networks for audio pattern recognition
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D. Plumbley. 2020 · 2020
Cited alongside, same era.
Learning Music Helps You Read: Using transfer to study linguistic structure in language models
Isabel Papadimitriou and Dan Jurafsky. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Cited alongside, same era.
RAVE: A variational autoencoder for fast and high-quality neural audio synthesis
Antoine Caillon and Philippe Esling. 2021 · 2021
Cited alongside, same era.
Unsupervised audiovisual synthesis via exemplar autoencoders
Kangle Deng, Aayush Bansal, and Deva Ramanan. 2021 · 2021
Cited alongside, same era.
BDDM: bilateral denoising diffusion models for fast and high-quality speech synthesis
Max W. Y. Lam, Jun Wang, Dan Su, and Dong Yu. 2022 · 2022
Later among the works it cites.
Autoregressive image generation using residual quantization
Doyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho, and Wook-Shin Han. 2022 · 2022
Later among the works it cites.
Binauralgrad: A two-stage conditional diffusion probabilistic model for binaural audio synthesis
Yichong Leng, Zehua Chen, Junliang Guo, Haohe Liu, Jiawei Chen, Xu Tan, Danilo P. Mandic, Lei He, Xiang-Yang Li, Tao Qin, Sheng Zhao, and Tie-Yan Liu. 2022 · 2022
Later among the works it cites.
Clip-event: Connecting text and images with event structures
Manling Li, Ruochen Xu, Shuohang Wang, Luowei Zhou, Xudong Lin, Chenguang Zhu, Michael Zeng, Heng Ji, and Shih-Fu Chang. 2022 · 2022
Later among the works it cites.
Chunked autoregressive GAN for conditional waveform synthesis
Max Morrison, Rithesh Kumar, Kundan Kumar, Prem Seetharaman, Aaron C. Courville, and Yoshua Bengio. 2022 · 2022
Later among the works it cites.
Musika! fast infinite waveform music generation
Marco Pasini and Jan Schlüter. 2022 · 2022
Later among the works it cites.
Diffusion autoencoders: Toward a meaningful and decodable representation
Konpat Preechakul, Nattanat Chatthee, Suttisak Wizadwongsa, and Supasorn Suwajanakorn. 2022 · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with CLIP latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022 · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Later among the works it cites.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. 2022 · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi. 2022 · 2022
Later among the works it cites.
Progressive distillation for fast sampling of diffusion models
Tim Salimans and Jonathan Ho. 2022 · 2022
Later among the works it cites.
Phenaki: Variable length video generation from open domain textual description
Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans, Hernan Moraldo, Han Zhang, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, and Dumitru Erhan. 2022 · 2022
Later among the works it cites.
Diffsound: Discrete diffusion model for text-to-sound generation
Dongchao Yang, Jianwei Yu, Helin Wang, Wen Wang, Chao Weng, Yuexian Zou, and Dong Yu. 2022 · 2022
Later among the works it cites.
Muse: Text-to-image generation via masked generative transformers
Huiwen Chang, Han Zhang, Jarred Barber, Aaron Maschinot, José Lezama, Lu Jiang, Ming-Hsuan Yang, Kevin Murphy, William T. Freeman, Michael Rubinstein, Yuanzhen Li, and Dilip Krishnan. 2023 · 2023
Closest in time.
Data rivers: Carving out the public domain in the age of Chat-GPT
Sylvie Delacroix. 2023 · 2023
Closest in time.
ArchiSound: Audio generation with diffusion
Flavio Schneider. 2023 · 2023
Closest in time.
Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. 2023 · 2023
Closest in time.