Fetching the paper…
Reading the bibliography…
Video-to-music generation demands both a temporally localized high-quality listening experience and globally aligned video-acoustic signatures.
Encoding musical style with transformer autoencoders
Choi, K.; Hawthorne, C.; Simon, I.; Dinculescu, M.; and Engel, J. 2019 · 1912
Earlier work this paper cites.
Towards learning a universal non-semantic representation of speech
Shor, J.; Jansen, A.; Maor, R.; Lang, O.; Tuval, O.; Quitry, F. d. C.; Tagliasacchi, M.; Shavitt, I.; Emanuel, D.; and Haviv, Y. 2020 · 2002
Earlier work this paper cites.
Jukebox: A generative model for music
Dhariwal, P.; Jun, H.; Payne, C.; Kim, J. W.; Radford, A.; and Sutskever, I. 2020 · 2005
Earlier work this paper cites.
Multi-Instrumentalist Net: Unsupervised Generation of Music from Body Movements
Su, K.; Liu, X.; and Shlizerman, E. 2020b · 2012
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Tran, D.; Bourdev, L.; Fergus, R.; Torresani, L.; and Paluri, M. 2015 · 2015
Earlier work this paper cites.
Youtube-8M: A large-scale video classification benchmark
Abu-El-Haija, S.; Kothari, N.; Lee, J.; Natsev, P.; Toderici, G.; Varadarajan, B.; and Vijayanarasimhan, S. 2016 · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Oord, A. v. d.; et al. 2016 · 2016
Earlier work this paper cites.
Visually indicated sounds
Owens, A.; Isola, P.; McDermott, J.; Torralba, A.; Adelson, E. H.; and Freeman, W. T. 2016 · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J.; and Zisserman, A. 2017 · 2017
Earlier work this paper cites.
Deep cross-modal audio-visual generation
Chen, L.; Srivastava, S.; Duan, Z.; and Xu, C. 2017 · 2017
Earlier work this paper cites.
Neural audio synthesis of musical notes with wavenet autoencoders
Engel, J.; et al. 2017 · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F.; Ellis, D. P.; Freedman, D.; Jansen, A.; Lawrence, W.; Moore, R. C.; Plakal, M.; and Ritter, M. 2017 · 2017
Earlier work this paper cites.
CNN architectures for large-scale audio classification
Hershey, S.; Chaudhuri, S.; Ellis, D. P.; Gemmeke, J. F.; Jansen, A.; et al. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Visually indicated sound generation by perceptually optimized classification
Chen, K.; Zhang, C.; Fang, C.; Wang, Z.; Bui, T.; and Nevatia, R. 2018 · 2018
Earlier work this paper cites.
Musegan: Multi-track sequential generative adversarial networks for symbolic music generation and accompaniment
Dong, H.-W.; Hsiao, W.-Y.; Yang, L.-C.; and Yang, Y.-H. 2018 · 2018
Earlier work this paper cites.
Enabling factorized piano music modeling and generation with the MAESTRO dataset
Hawthorne, C.; et al. 2018 · 2018
Earlier work this paper cites.
Huang, C.-Z.; et al. 2018 · 2018
Earlier work this paper cites.
The 2nd youtube-8M large-scale video understanding challenge
Lee, J.; Reade, W.; Sukthankar, R.; Toderici, G.; et al. 2018 · 2018
Cited alongside, same era.
A hierarchical latent vector model for learning long-term structure in music
Roberts, A.; Engel, J.; Raffel, C.; Hawthorne, C.; and Eck, D. 2018 · 2018
Cited alongside, same era.
Visual to sound: Generating natural sound for videos in the wild
Zhou, Y.; Wang, Z.; Fang, C.; Bui, T.; and Berg, T. L. 2018 · 2018
Cited alongside, same era.
Learning to Groove with Inverse Sequence Transformations
Gillick, J.; Roberts, A.; Engel, J.; Eck, D.; and Bamman, D. 2019 · 2019
Cited alongside, same era.
High-level control of drum track generation using learned patterns of rhythmic interaction
Lattner, S.; and Grachten, M. 2019 · 2019
Cited alongside, same era.
Generating Visually Aligned Sound from Videos
Chen, P.; Zhang, Y.; Tan, M.; Xiao, H.; Huang, D.; and Gan, C. 2020 · 2020
Vector-quantized image modeling with improved VQGAN
Yu, J.; Li, X.; Koh, J. Y.; Zhang, H.; Pang, R.; Qin, J.; Ku, A.; Xu, Y.; Baldridge, J.; and Wu, Y. 2021 · 2021
Later among the works it cites.
AudioLM: a language modeling approach to audio generation
Borsos, Z.; Marinier, R.; Vincent, D.; Kharitonov, E.; Pietquin, O.; Sharifi, M.; et al. 2022 · 2022
Later among the works it cites.
High Fidelity Neural Audio Compression
D’efossez, A.; et al. 2022 · 2022
Later among the works it cites.
Riffusion - Stable diffusion for real-time music generation
Forsgren, S.; and Martiros, H. 2022 · 2022
Later among the works it cites.
MuLan: A joint embedding of music audio and natural language
Huang, Q.; Jansen, A.; Lee, J.; Ganti, R.; Li, J. Y.; and Ellis, D. P. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
X3D: Expanding architectures for efficient video recognition
Feichtenhofer, C. 2020 · 2020
Cited alongside, same era.
Foley Music: Learning to Generate Music from Videos
Gan, C.; Huang, D.; Chen, P.; Tenenbaum, J. B.; and Torralba, A. 2020 · 2020
Cited alongside, same era.
Pop Music Transformer: Beat-based modeling and generation of expressive Pop piano compositions
Huang, Y.-S.; and Yang, Y.-H. 2020 · 2020
Cited alongside, same era.
Sight to Sound: An End-to-End Approach for Visual Piano Transcription
Koepke, A. S.; Wiles, O.; Moses, Y.; and Zisserman, A. 2020 · 2020
Cited alongside, same era.
This time with feeling: Learning expressive musical performance
Oore, S.; Simon, I.; Dieleman, S.; Eck, D.; and Simonyan, K. 2020 · 2020
Cited alongside, same era.
ViViT: A video vision transformer
Arnab, A.; Dehghani, M.; Heigold, G.; Sun, C.; Lučić, M.; and Schmid, C. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
Audiogen: Textually guided audio generation
Kreuk, F.; Synnaeve, G.; Polyak, A.; Singer, U.; Défossez, A.; Copet, J.; Parikh, D.; Taigman, Y.; and Adi, Y. 2022 · 2022
Later among the works it cites.
It’s Time for Artistic Correspondence in Music and Video
Surís, D.; Vondrick, C.; Russell, B.; and Salamon, J. 2022 · 2022
Later among the works it cites.
Diffsound: Discrete diffusion model for text-to-sound generation
Yang, D.; Yu, J.; Wang, H.; Wang, W.; Weng, C.; Zou, Y.; and Yu, D. 2022 · 2022
Later among the works it cites.
MusicLM: Generating Music From Text
Agostinelli, A.; Denk, T. I.; Borsos, Z.; Engel, J.; Verzetti, M.; Caillon, A.; Huang, Q.; Jansen, A.; et al. 2023 · 2023
Closest in time.
Simple and Controllable Music Generation
Copet, J.; Kreuk, F.; Gat, I.; Remez, T.; Kant, D.; Synnaeve, G.; Adi, Y.; and Défossez, A. 2023 · 2023
Closest in time.
SingSong: Generating musical accompaniments from singing
Donahue, C.; et al. 2023 · 2023
Closest in time.
VampNet: Music Generation via Masked Acoustic Token Modeling
Garcia, H. F.; et al. 2023 · 2023
Closest in time.
High-Fidelity Audio Compression with Improved RVQGAN
Kumar, R.; et al. 2023 · 2023
Closest in time.
AudioLDM: Text-to-audio generation with latent diffusion models
Liu, H.; Chen, Z.; Yuan, Y.; Mei, X.; Liu, X.; Mandic, D.; Wang, W.; and Plumbley, M. D. 2023 · 2023
Closest in time.
Moûsai: Text-to-Music Generation with Long-Context Latent Diffusion
Schneider, F.; et al. 2023 · 2023
Closest in time.
Physics-Driven Diffusion Models for Impact Sound Synthesis from Videos
Su, K.; Qian, K.; Shlizerman, E.; Torralba, A.; and Gan, C. 2023 · 2023
Closest in time.
Long-Term Rhythmic Video Soundtracker
Yu, J.; Wang, Y.; Chen, X.; Sun, X.; and Qiao, Y. 2023 · 2023
Closest in time.