Fetching the paper…
Reading the bibliography…
Numerous studies in the field of music generation have demonstrated impressive performance, yet virtually no models are able to directly generate music to match accompanying videos.
A bi-directional transformer for musical chord recognition
Park, J., Choi, K., Jeon, S., Kim, D., & Park, J. (2019) · 1907
Earlier work this paper cites.
Music: an appreciation
Kamien, R., & Kamien, A. (1988) · 1988
Earlier work this paper cites.
Unheard melodies: Narrative film music
Littlefield, R. (1990) · 1990
Earlier work this paper cites.
Automatic background music generation based on actors’ mood and motions
Nakamura, J.-I., Kaku, T., Hyun, K., Noma, T., & Yoshida, S. (1994) · 1994
Earlier work this paper cites.
Magenta: An architecture for real time automatic composition of background music
Casella, P., & Paiva, A. (2001) · 2001
Earlier work this paper cites.
Cognitive foundations of musical pitch
Krumhansl, C. L. (2001) · 2001
Earlier work this paper cites.
Sound synthesis from real-time video images,
Dannenberg, R. B., & Neuendorffer, T. (2003) · 2003
Earlier work this paper cites.
About the determination of key of a musical excerpt
Bellmann, H. (2006) · 2005
Earlier work this paper cites.
How music really works!: the essential handbook for songwriters, performers, and music students
Chase, W. (2006) · 2006
Earlier work this paper cites.
The long zoom
Johnson, S. (2006) · 2006
Earlier work this paper cites.
Quantitative and visual analysis of the impact of music on perceived emotion of film
Parke, R., Chew, E., & Kyriakakis, C. (2007) · 2007
Earlier work this paper cites.
Music and probability
Temperley, D. (2007) · 2007
Earlier work this paper cites.
Game sound: an introduction to the history, theory, and practice of video game music and sound design
Collins, K. (2008) · 2008
Earlier work this paper cites.
Wu, S.-L., & Yang, Y.-H. (2020) · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., & Fei-Fei, L. (2009) · 2009
Earlier work this paper cites.
music21: A toolkit for computer-aided musicology and symbolic music data,
Cuthbert, M. S., & Ariza, C. (2010) · 2010
Earlier work this paper cites.
Determination of nonprototypical valence and arousal in popular music: features and performances
Schuller, B., Dorfner, J., & Rigoll, G. (2010) · 2010
Earlier work this paper cites.
Experience-driven procedural music generation for games
Plans, D., & Morelli, D. (2012) · 2012
Earlier work this paper cites.
Multi-instrumentalist net: Unsupervised generation of music from body movements
Su, K., Liu, X., & Shlizerman, E. (2020b) · 2012
Earlier work this paper cites.
Composing fifth species counterpoint music with a variable neighborhood search algorithm
Herremans, D., & Sörensen, K. (2013) · 2013
Earlier work this paper cites.
Polyphonic music generation by modeling temporal dependencies using a rnn-dbn
Goel, K., Vohra, R., & Sahoo, J. K. (2014) · 2014
Earlier work this paper cites.
Mir_eval: A transparent implementation of common mir metrics
Raffel, C., McFee, B., Humphrey, E. J., Salamon, J., Nieto, O., Liang, D., Ellis, D. P., & Raffel, C. C. (2014) · 2014
Earlier work this paper cites.
Automatic real-time music generation for games
Engels, S., Tong, T., & Chan, F. (2015) · 2015
Earlier work this paper cites.
Generating structured music for bagana using quality metrics based on markov models
Herremans, D., Weisser, S., Sörensen, K., & Conklin, D. (2015) · 2015
Earlier work this paper cites.
Chordripple: Recommending chords to help novice composers go beyond the ordinary
Huang, C.-Z. A., Duvenaud, D., & Gajos, K. Z. (2016) · 2016
Cited alongside, same era.
On the potential of simple framewise approaches to piano transcription
Kelz, R., Dorfer, M., Korzeniowski, F., Böck, S., Arzt, A., & Widmer, G. (2016) · 2016
Cited alongside, same era.
Adaptive music generation for computer games
Prechtl, A. (2016) · 2016
Cited alongside, same era.
Morpheus: generating structured music with constrained patterns and tension
Herremans, D., & Chew, E. (2017) · 2017
Cited alongside, same era.
A functional taxonomy of music generation systems
Herremans, D., Chuan, C.-H., & Chew, E. (2017) · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017) · 2017
Vivit: A video vision transformer
Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lučić, M., & Schmid, C. (2021) · 2021
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., & Mordatch, I. (2021) · 2021
Later among the works it cites.
Reconvat: A semi-supervised automatic music transcription framework for low-resource real-world data
Cheuk, K. W., Herremans, D., & Su, L. (2021) · 2021
Later among the works it cites.
Chord conditioned melody generation with transformer based decoders
Choi, K., Park, J., Heo, W., Jeon, S., & Park, J. (2021) · 2021
Later among the works it cites.
Controllable deep melody generation via hierarchical music structure representation
Dai, S., Jin, Z., Gomes, C., & Dannenberg, R. B. (2021) · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Spatial-aware hierarchical collaborative deep learning for poi recommendation
Yin, H., Wang, W., Wang, H., Chen, L., & Zhou, X. (2017) · 2017
Cited alongside, same era.
Pyscenedetect: Intelligent scene cut detection and video splitting tool
Castellano, B. (2018) · 2018
Cited alongside, same era.
Modeling temporal tonal relations in polyphonic music through deep networks with a novel image-based representation
Chuan, C.-H., & Herremans, D. (2018) · 2018
Cited alongside, same era.
Huang, C.-Z. A., Vaswani, A., Uszkoreit, J., Shazeer, N., Simon, I., Hawthorne, C., Dai, A. M., Hoffman, M. D., Dinculescu, M., & Eck, D. (2018) · 2018
Cited alongside, same era.
Self-attention with relative position representations
Shaw, P., Uszkoreit, J., & Vaswani, A. (2018) · 2018
Cited alongside, same era.
Neural memory streaming recommender networks with adversarial training
Wang, Q., Yin, H., Hu, Z., Lian, D., Wang, H., & Huang, Z. (2018) · 2018
Cited alongside, same era.
Gong, Y., Chung, Y.-A., & Glass, J. (2021) · 2021
Later among the works it cites.
Generating lead sheets with affect: A novel conditional seq2seq framework
Makris, D., Agres, K. R., & Herremans, D. (2021) · 2021
Later among the works it cites.
Symbolic music generation with diffusion models
Mittal, G., Engel, J., Hawthorne, C., & Simon, I. (2021) · 2021
Later among the works it cites.
Symbolic music generation with transformer-gans
Muhamed, A., Li, L., Shi, X., Yaddanapudi, S., Chi, W., Jackson, D., Suresh, R., Lipton, Z. C., & Smola, A. J. (2021) · 2021
Later among the works it cites.
Deep-learning-based multimodal emotion classification for music videos
Pandeya, Y. R., Bhattarai, B., & Lee, J. (2021) · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J. et al. (2021) · 2021
Later among the works it cites.
How does it sound?
Su, K., Liu, X., & Shlizerman, E. (2021) · 2021
Later among the works it cites.
Calliope–a polyphonic shaw2018selfsformer
Valenti, A., Berti, S., & Bacciu, D. (2021) · 2021
Later among the works it cites.
Omnizart: A general toolbox for automatic shaw2018selfscription
Wu, Y.-T., Luo, Y.-J., Chen, T.-P., Wei, I.-C., Hsu, J.-Y., Chuang, Y.-C., & Su, L. (2021) · 2021
Later among the works it cites.
Musicbert: Symbolic music understanding with large-scale pre-training
Zeng, M., Tan, X., Wang, R., Ju, Z., Qin, T., & Liu, T.-Y. (2021) · 2021
Later among the works it cites.
Hierarchical recurrent neural networks for conditional melody generation with long-term structure
Zixun, G., Makris, D., & Herremans, D. (2021) · 2021
Later among the works it cites.
A systematic review of artificial intelligence-based music generation: Scope, applications, and future trends
Civit, M., Civit-Masot, J., Cuadrado, F., & Escalona, M. J. (2022) · 2022
Later among the works it cites.
Conditional drums generation using compound word representations
Makris, D., Zixun, G., Kaliakatsos-Papakostas, M., & Herremans, D. (2022) · 2022
Later among the works it cites.
Quantized gan for complex music generation from dance videos
Zhu, Y., Olszewski, K., Wu, Y., Achlioptas, P., Chai, M., Yan, Y., & Tulyakov, S. (2022a) · 2022
Later among the works it cites.
Diffroll: Diffusion-based generative music transcription with unsupervised pretraining capability
Cheuk, K. W., Sawata, R., Uesaka, T., Murata, N., Takahashi, N., Takahashi, S., Herremans, D., & Mitsufuji, Y. (2023) · 2023
Closest in time.
V2meow: Meowing to the visual beat via music generation
Su, K., Li, J. Y., Huang, Q., Kuzmin, D., Lee, J., Donahue, C., Sha, F., Jansen, A., Wang, Y., Verzetti, M. et al. (2023) · 2023
Closest in time.
Emomv: Affective music-video correspondence learning datasets for classification and retrieval
Thao, H. T. P., Roig, G., & Herremans, D. (2023) · 2023
Closest in time.
Video background music generation with controllable shaw2018selfsformer
Di, S., Jiang, Z., Liu, S., Wang, Z., Zhu, L., He, Z., Liu, H., & Yan, S. (2021) · 2045
Closest in time.