Fetching the paper…
Reading the bibliography…
We introduce UniMuMo, a unified multimodal model capable of taking arbitrary text, music, and motion data as input conditions to generate outputs across all three modalities.
Fundamentals of speech recognition
Lawrence R Rabiner and Biing-Hwang Juang · 1999
Earlier work this paper cites.
On synchronizing movements to music
Edward W Large · 2000
Earlier work this paper cites.
Beat tracking by dynamic programming
Daniel PW Ellis · 2007
Earlier work this paper cites.
Computing and visualizing dynamic time warping alignments in r: the dtw package
Toni Giorgino · 2009
Earlier work this paper cites.
Musical movement and synchronization
Peter E Keller and Martina Rieger · 2009
Earlier work this paper cites.
Maximum filter vibrato suppression for onset detection
Sebastian Böck and Gerhard Widmer · 2013
Earlier work this paper cites.
librosa: Audio and music signal analysis in python
Brian McFee, Colin Raffel, Dawen Liang, Daniel PW Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto · 2015
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Visual rhythm and beat
Abe Davis and Maneesh Agrawala · 2018
Earlier work this paper cites.
Learning a bidirectional mapping between human whole-body motion and natural language using deep recurrent neural networks
Matthias Plappert, Christian Mandery, and Tamim Asfour · 2018
Earlier work this paper cites.
Paired recurrent autoencoders for bidirectional translation between robot actions and linguistic descriptions
Tatsuro Yamada, Hiroyuki Matsunaga, and Tetsuya Ogata · 2018
Earlier work this paper cites.
AMASS: Archive of motion capture as surface shapes
Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll, and Michael J. Black · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Foley music: Learning to generate music from videos
Chuang Gan, Deng Huang, Peihao Chen, Joshua B Tenenbaum, and Antonio Torralba · 2020
Earlier work this paper cites.
Music4all: A new music database and its applications
Igor André Pegoraro Santana, Fabio Pinhelli, Juliano Donini, Leonardo Catharin, Rafael Biazus Mangolin, Valéria Delisandra Feltrim, Marcos Aurélio Domingues, et al · 2020
Earlier work this paper cites.
Dance2music: Automatic dance-driven music generation
Gunjan Aggarwal and Devi Parikh · 2021
Cited alongside, same era.
Drum-aware ensemble architecture for improved joint musical beat and downbeat tracking
Ching-Yu Chiu, Alvin Wen-Yu Su, and Yi-Hsuan Yang · 2021
Cited alongside, same era.
Hybrid spectrogram and waveform source separation
Alexandre Défossez · 2021
Cited alongside, same era.
Video background music generation with controllable music transformer
Shangzhe Di, Zeren Jiang, Si Liu, Zhaokai Wang, Leyan Zhu, Zexin He, Hongming Liu, and Shuicheng Yan · 2021
Cited alongside, same era.
Linguistic descriptions of human motion with generative adversarial seq2seq learning
Yusuke Goutsu and Tetsunari Inamura · 2021
Cited alongside, same era.
How does it sound?
Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Daniel Cohen-Or, and Amit H Bermano · 2022
Later among the works it cites.
Motiondiffuse: Text-driven human motion generation with diffusion model
Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, and Ziwei Liu · 2022
Later among the works it cites.
Musiclm: Generating music from text
Andrea Agostinelli, Timo I. Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, Matthew Sharifi, Neil Zeghidour, and C. Frank · 2023
Later among the works it cites.
Executing your commands via motion diffusion in latent space
Xin Chen, Biao Jiang, Wen Liu, Zilong Huang, Bin Fu, Tao Chen, and Gang Yu · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kun Su, Xiulong Liu, and Eli Shlizerman · 2021
Cited alongside, same era.
High fidelity neural audio compression
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi · 2022
Cited alongside, same era.
A bi-directional attention guided cross-modal network for music based dance generation
Di Fan, Lili Wan, Wanru Xu, and Shenghui Wang · 2022
Cited alongside, same era.
Riffusion - Stable diffusion for real-time music generation
Seth* Forsgren and Hayk* Martiros · 2022
Cited alongside, same era.
Audiogen: Textually guided audio generation
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi · 2022
Cited alongside, same era.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Cited alongside, same era.
Mubert. https://mubert. com/, https://github.com/mubertai/ mubert-text-to-music
Mubert-Inc · 2022
Cited alongside, same era.
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez · 2023
Later among the works it cites.
Lp-musiccaps: Llm-based pseudo music captioning
SeungHeon Doh, Keunwoo Choi, Jongpil Lee, and Juhan Nam · 2023
Later among the works it cites.
Yuan Gong, Hongyin Luo, Alexander H Liu, Leonid Karlinsky, and James Glass · 2023
Later among the works it cites.
Motiongpt: Human motion as a foreign language
Biao Jiang, Xin Chen, Wen Liu, Jingyi Yu, Gang Yu, and Tao Chen · 2023
Later among the works it cites.
A whisper transformer for audio captioning trained with synthetic captions and transfer learning
Marek Kadlčík, Adam Hájek, Jürgen Kieslich, and Radosław Winiecki · 2023
Later among the works it cites.
Hybrid transformers for music source separation
Simon Rouard, Francisco Massa, and Alexandre Défossez · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models. corr, abs/2302.13971, 2023. doi: 10.48550
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Edge: Editable dance generation from music
Jonathan Tseng, Rodrigo Castellon, and Karen Liu · 2023
Later among the works it cites.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov · 2023
Later among the works it cites.
T2m-gpt: Generating human motion from textual descriptions with discrete representations
Jianrong Zhang, Yangsong Zhang, Xiaodong Cun, Shaoli Huang, Yong Zhang, Hongwei Zhao, Hongtao Lu, and Xi Shen · 2023
Later among the works it cites.