Fetching the paper…
Reading the bibliography…
CLIP (Contrastive Language-Image Pre-Training) is a multimodal neural network trained on (text, image) pairs to predict the most relevant text caption given an image.
The million song dataset
Thierry Bertin-Mahieux, Daniel P.W. Ellis, Brian Whitman, and Paul Lamere · 2011
Earlier work this paper cites.
Mumu: Multimodal music dataset, July 2017
Sergio Oramas · 2015
Earlier work this paper cites.
Generative adversarial text to image synthesis
Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee · 2016
Earlier work this paper cites.
Learning-based methods for comparing sequences, with applications to audio-to-midi alignment and matching, July 2016
Colin Raffel · 2016
Earlier work this paper cites.
Musegan: Multi-track sequential generative adversarial networks for symbolic music generation and accompaniment, 2017
Hao-Wen Dong, Wen-Yi Hsiao, Li-Chia Yang, and Yi-Hsuan Yang · 2017
Earlier work this paper cites.
Onsets and frames: Dual-objective piano transcription, 2017
Curtis Hawthorne, Erich Elsen, Jialin Song, Adam Roberts, Ian Simon, Colin Raffel, Jesse Engel, Sageev Oore, and Douglas Eck · 2017
Earlier work this paper cites.
A hierarchical latent vector model for learning long-term structure in music
Adam Roberts, Jesse Engel, Colin Raffel, Curtis Hawthorne, and Douglas Eck · 2018
Earlier work this paper cites.
An improved relative self-attention mechanism for transformer with application to music generation
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Noam Shazeer, Curtis Hawthorne, Andrew M. Dai, Matthew D. Hoffman, and Douglas Eck · 2018
Cited alongside, same era.
Large scale GAN training for high fidelity natural image synthesis
Andrew Brock, Jeff Donahue, and Karen Simonyan · 2018
Cited alongside, same era.
Melgan: Generative adversarial networks for conditional waveform synthesis
Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brébisson, Yoshua Bengio, and Aaron C Courville · 2019
Cited alongside, same era.
Pop Music Transformer: Beat-Based Modeling and Generation of Expressive Pop Piano Compositions
Yu-Siang Huang and Yi-Hsuan Yang · 2020
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Musemorphose: Full-song and fine-grained music style transfer with just one transformer VAE
Shih-Lun Wu and Yi-Hsuan Yang · 2021
Later among the works it cites.
Clip-guided gan image generation: An artistic exploration
Amy Smith and Simon Colton · 2021
Later among the works it cites.
Compound word transformer: Learning to compose full-song music over dynamic directed hypergraphs, 2021
Wen-Yi Hsiao, Jen-Yu Liu, Yin-Cheng Yeh, and Yi-Hsuan Yang · 2021
Later among the works it cites.
Musemorphose: Full-song and fine-grained music style transfer with one transformer vae, 2021
Shih-Lun Wu and Yi-Hsuan Yang · 2021
Later among the works it cites.
Theme transformer: Symbolic music generation with theme-conditioned transformer, 2021
Yi-Jen Shih, Shih-Lun Wu, Frank Zalkow, Meinard Müller, and Yi-Hsuan Yang · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Cited alongside, same era.
Theme transformer: Symbolic music generation with theme-conditioned transformer, 2021
Yi-Jen Shih, Shih-Lun Wu, Frank Zalkow, Meinard Müller, and Yi-Hsuan Yang · 2021
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents, 2022
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Later among the works it cites.