Fetching the paper…
Reading the bibliography…
The seamless integration of music with dance movements is essential for communicating the artistic intent of a dance piece.
Automatic synchronization of background music and motion in computer animation. In Computer Graphics Forum , Vol. 24. Amsterdam: North Holland, 1982-, 353–362
Hyun-Chul Lee, In-Kwon Lee, et al · 1982
Earlier work this paper cites.
YouTube-8M: A Large-Scale Video Classification Benchmark
Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan. 2016 · 2016
Earlier work this paper cites.
CNN architectures for large-scale audio classification. In International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 131–135
Shawn Hershey, Sourish Chaudhuri, Daniel PW Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al · 2017
Earlier work this paper cites.
Visual rhythm and beat
Abe Davis and Maneesh Agrawala. 2018 · 2018
Earlier work this paper cites.
Fréechet Audio Distance: A Reference-free Metric for Evaluating Music Enhancement Algorithms. In INTERSPEECH . 2350–2354
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi. 2019 · 2019
Earlier work this paper cites.
Deep high-resolution representation learning for human pose estimation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 5693–5703
Ke Sun, Bin Xiao, Dong Liu, and Jingdong Wang. 2019 · 2019
Earlier work this paper cites.
OpenMMLab Pose Estimation Toolbox and Benchmark
MMPose Contributors. 2020 · 2020
Earlier work this paper cites.
Foley music: Learning to generate music from videos. In European Conference on Computer Vision (ECCV) . Springer, 758–775
Chuang Gan, Deng Huang, Peihao Chen, Joshua B Tenenbaum, and Antonio Torralba. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Earlier work this paper cites.
Audeo: Audio generation for a silent performance video
Kun Su, Xiulong Liu, and Eli Shlizerman. 2020 · 2020
Earlier work this paper cites.
Dance2Music: Automatic Dance-driven Music Generation
Gunjan Aggarwal and Devi Parikh. 2021 · 2021
Earlier work this paper cites.
LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations (ICLR)
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021 · 2021
Earlier work this paper cites.
AI Choreographer: Music Conditioned 3D Dance Generation with AIST++. In IEEE/CVF International Conference on Computer Vision (ICCV) . 13381–13392
Ruilong Li, Shan Yang, David A. Ross, and Angjoo Kanazawa. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision. In International Conference on Machine Learning (ICML) . PMLR, 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
How Does it Sound? Generation of Rhythmic Soundtracks for Human Movement Videos. In Thirty-fifth Conference on Neural Information Processing Systems (NeurIPS)
Kun Su, Xiulong Liu, and Eli Shlizerman. 2021 · 2021
Earlier work this paper cites.
Riffusion - Stable diffusion for real-time music generation
Seth* Forsgren and Hayk* Martiros. 2022 · 2022
Earlier work this paper cites.
NeuralSound: learning-based modal sound synthesis with acoustic transfer
Xutong Jin, Sheng Li, Guoping Wang, and Dinesh Manocha. 2022 · 2022
Earlier work this paper cites.
High-Resolution Image Synthesis with Latent Diffusion Models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 10684–10695
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Earlier work this paper cites.
MusicLM: Generating Music From Text
Andrea Agostinelli, Timo I Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, et al · 2023
Earlier work this paper cites.
A Neural Space-Time Representation for Text-to-Image Personalization
Yuval Alaluf, Elad Richardson, Gal Metzer, and Daniel Cohen-Or. 2023 · 2023
Cited alongside, same era.
Simple and Controllable Music Generation. In Thirty-seventh Conference on Neural Information Processing Systems (NeurIPS)
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez. 2023 · 2023
Cited alongside, same era.
Conditional Generation of Audio from Video via Foley Analogies. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 2426–2436
Yuexi Du, Ziyang Chen, Justin Salamon, Bryan Russell, and Andrew Owens. 2023 · 2023
Cited alongside, same era.
Encoder-Based Domain Tuning for Fast Personalization of Text-to-Image Models
Rinon Gal, Moab Arar, Yuval Atzmon, Amit H. Bermano, Gal Chechik, and Daniel Cohen-Or. 2023b · 2023
Cited alongside, same era.
LLark: A Multimodal Foundation Model for Music
Josh Gardner, Simon Durand, Daniel Stoller, and Rachel M Bittner. 2023 · 2023
Any-to-Any Generation via Composable Diffusion
Zineng Tang, Ziyi Yang, Chenguang Zhu, Michael Zeng, and Mohit Bansal. 2023 · 2023
Later among the works it cites.
P + P+ : Extended Textual Conditioning in Text-to-Image Generation
Andrey Voynov, Qinghao Chu, Daniel Cohen-Or, and Kfir Aberman. 2023 · 2023
Later among the works it cites.
Edict: Exact diffusion inversion via coupled transformations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 22532–22541
Bram Wallace, Akash Gokul, and Nikhil Naik. 2023 · 2023
Later among the works it cites.
NExT-GPT: Any-to-Any Multimodal LLM
Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. 2023c · 2023
Later among the works it cites.
Music ControlNet: Multiple Time-varying Controls for Music Generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Text-to-Audio Generation using Instruction Guided Latent Diffusion Model. In ACM International Conference on Multimedia (Ottawa ON, Canada). Association for Computing Machinery, New York, NY, USA, 3590–3598
Deepanway Ghosal, Navonil Majumder, Ambuj Mehrish, and Soujanya Poria. 2023 · 2023
Cited alongside, same era.
Noise2Music: Text-conditioned Music Generation with Diffusion Models
Qingqing Huang, Daniel S Park, Tao Wang, Timo I Denk, Andy Ly, Nanxin Chen, Zhengdong Zhang, Zhishuai Zhang, Jiahui Yu, Christian Frank, et al · 2023
Cited alongside, same era.
ReVersion: Diffusion-Based Relation Inversion from Images
Ziqi Huang, Tianxing Wu, Yuming Jiang, Kelvin CK Chan, and Ziwei Liu. 2023c · 2023
Cited alongside, same era.
StyleDiffusion: Prompt-Embedding Inversion for Text-Based Editing
Senmao Li, Joost van de Weijer, Taihang Hu, Fahad Shahbaz Khan, Qibin Hou, Yaxing Wang, and Jian Yang. 2023 · 2023
Cited alongside, same era.
AudioLDM: Text-to-Audio Generation with Latent Diffusion Models. In International Conference on Machine Learning (ICML) , Vol. 202. PMLR, 21450–21474
Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, and Mark D Plumbley. 2023 · 2023
Cited alongside, same era.
MuseCoco: Generating Symbolic Music from Text
Peiling Lu, Xin Xu, Chenfei Kang, Botao Yu, Chengyi Xing, Xu Tan, and Jiang Bian. 2023 · 2023
Cited alongside, same era.
Null-text Inversion for Editing Real Images using Guided Diffusion Models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 6038–6047
Ron Mokady, Amir Hertz, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. 2023 · 2023
Cited alongside, same era.
Shih-Lun Wu, Chris Donahue, Shinji Watanabe, and Nicholas J Bryan. 2023b · 2023
Later among the works it cites.
Long-Term Rhythmic Video Soundtracker. In International Conference on Machine Learning (ICML) (Honolulu, Hawaii, USA). JMLR.org, Article 1688, 15 pages
Jiashuo Yu, Yaohui Wang, Xinyuan Chen, Xiao Sun, and Yu Qiao. 2023 · 2023
Later among the works it cites.
ProSpect: Prompt Spectrum for Attribute-Aware Personalization of Diffusion Models
Yuxin Zhang, Weiming Dong, Fan Tang, Nisha Huang, Haibin Huang, Chongyang Ma, Tong-Yee Lee, Oliver Deussen, and Changsheng Xu. 2023a · 2023
Later among the works it cites.
Video background music generation: Dataset, method and evaluation. In IEEE/CVF International Conference on Computer Vision (ICCV) . 15637–15647
Le Zhuo, Zhaokai Wang, Baisen Wang, Yue Liao, Chenxi Bao, Stanley Peng, Songhao Han, Aixi Zhang, Fei Fang, and Si Liu. 2023 · 2023
Later among the works it cites.
Dance2MIDI: Dance-driven multi-instrument music generation
Bo Han, Yuheng Li, Yixuan Shen, Yi Ren, and Feilin Han. 2024 · 2024
Closest in time.
Video2Music: Suitable music generation from videos using an Affective Multimodal Transformer model
Jaeyong Kang, Soujanya Poria, and Dorien Herremans. 2024 · 2024
Closest in time.
Music Style Transfer with Time-Varying Inversion of Diffusion Models. In The Thirty-Eighth AAAI Conference on Artificial Intelligence (AAAI) , Vol. 38. 547–555
Sifei Li, Yuxin Zhang, Fan Tang, Chongyang Ma, Weiming Dong, and Changsheng Xu. 2024 · 2024
Closest in time.
AudioLDM 2: Learning Holistic Audio Generation With Self-Supervised Pretraining
Haohe Liu, Yi Yuan, Xubo Liu, Xinhao Mei, Qiuqiang Kong, Qiao Tian, Yuping Wang, Wenwu Wang, Yuxuan Wang, and Mark D. Plumbley. 2024 · 2024
Closest in time.
Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion. In International Conference on Machine Learning (ICML)
Hila Manor and Tomer Michaeli. 2024 · 2024
Closest in time.
Mustango: Toward controllable text-to-music generation
Jan Melechovsky, Zixun Guo, Deepanway Ghosal, Navonil Majumder, Dorien Herremans, and Soujanya Poria. 2024 · 2024
Closest in time.
DITTO: Diffusion Inference-Time T-Optimization for Music Generation
Zachary Novack, Julian McAuley, Taylor Berg-Kirkpatrick, and Nicholas J. Bryan. 2024 · 2024
Closest in time.
Investigating Personalization Methods in Text to Music Generation. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 1081–1085
Manos Plitsis, Theodoros Kouzelis, Georgios Paraskevopoulos, Vassilis Katsouros, and Yannis Panagakis. 2024 · 2024
Closest in time.
V2Meow: Meowing to the Visual Beat via Music Generation. In The Thirty-Eighth AAAI Conference on Artificial Intelligence (AAAI) . 4952–4960
Kun Su, Judith Yue Li, Qingqing Huang, Dima Kuzmin, Joonseok Lee, Chris Donahue, Fei Sha, Aren Jansen, Yu Wang, Mauro Verzetti, et al · 2024
Closest in time.
Video background music generation with controllable music transformer. In ACM International Conference on Multimedia . 2037–2045
Shangzhe Di, Zeren Jiang, Si Liu, Zhaokai Wang, Leyan Zhu, Zexin He, Hongming Liu, and Shuicheng Yan. 2021 · 2045
Closest in time.